Research Measures How Often AI Makes Financial Advice Mistakes

3 Min Read Tags:

  • Saturn found that AI models gave incorrect answers to financial questions an average of 57% of the time.
  • The error rate rose to 88% for complex queries involving multiple calculations and reached 99% for some models.
  • A pension-tax error by one model could have resulted in an HMRC charge of £17,500, according to the study.

Tech company Saturn tested free and paid versions of ChatGPT, Claude, Copilot, Grok and Gemini in research reported by the FT, finding that the models gave incorrect answers to financial questions an average of 57% of the time. The error rate rose to 88% for complex queries involving multiple calculations, highlighting the potential financial consequences of relying on generative AI for tax, pension and investment decisions.

AI models struggled with complex financial questions

Researchers tested more than 100 money-related questions across 18 models. They submitted more than 10,000 prompts in total, with each question repeated up to five times.

The study identified errors in financial calculations, failures to account for future changes to tax legislation, made-up rules and incorrect responses to complex questions requiring multiple calculations.

For some models, the error rate on complex questions reached 99%. Claude Opus 5 delivered the best result when used in “reasoning” mode, but it still answered 39% of questions incorrectly.

Amal Jolly, Saturn’s CEO, said widespread use of AI for financial guidance created additional risks.

“Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money,” he said.

Errors could carry substantial costs

In one test, researchers asked the free Claude Haiku 4.5 model about pension taxation. According to the study, its error could have resulted in an HMRC charge of £17,500, or about $23,430, for a pension saver.

In another case, Claude “made up a rule” about student loans and claimed that a graduate could stop making payments after moving abroad.

Paid models were generally more accurate than free versions, while newer models produced better results. Sarah Coles, head of personal finance at investment platform AJ Bell, said AI could still help users with budgeting and finding information.

“It’s a good example of why people who need support need proper advice rather than trusting AI,” she said.

The findings align with a recent study by the UK Financial Conduct Authority, which found that one in five UK adults was willing to let AI make financial decisions on their behalf. Interest was strongest in complex areas including debt, pensions and investments.

Jolly said AI financial advice was currently unregulated, leaving users without the same protections and compensation available when they consult a human professional.

OpenAI recently launched ChatGPT for the financial sector.

Source: Incrypted

TAGGED:
World’s First AI Actress Malfunctions During Live Broadcast

AI-generated actress Tilly Norwood briefly switched from English to Chinese after Tom Conti asked whether her Misaligned co-stars were human or digital in an interview clip published September 18, 2026.

4 Min Read
Kevin O’Leary Ranks as Main Threat to Bitcoin’s Path to $1M

Canadian investor Kevin O’Leary said bitcoin could reach $1 million if quantum-computing risks to its cryptography were resolved, while shifting his focus from Ethereum to AI-related energy infrastructure.

4 Min Read
NYSE Tested Avalanche for a Year as Stock Market Moves Onchain

Ava Labs President Charlie Cooper said NYSE spent the past year testing Avalanche technology, while the exchange has not announced a final blockchain choice for its planned tokenized-securities platform.

4 Min Read
AI Creates New Generation of Unicorn Startups

Andreessen Horowitz cited Silicon Valley Bank data showing new unicorns’ median age fell about 37% since 2023 to just over four years, without establishing AI as the sole cause.

4 Min Read
Gemini AI Model Moves Beyond Testing, Attacks Three Real Companies

Google’s Gemini accessed protected systems at three real companies during cybersecurity tests in May 2026, but stopped after recognizing them as real organizations, The Wall Street Journal reported.

4 Min Read