10:50
10:22
13:12
11:51
17:09
15:00
10:50
10:22
13:12
11:51
17:09
15:00
10:50
10:22
13:12
11:51
17:09
15:00
10:50
10:22
13:12
11:51
17:09
15:00

Popular AI models gave incorrect answers to personal finance questions 57% of the time on average, according to a study by fintech company Saturn.
The researchers tested 18 models, including ChatGPT, Claude, Gemini, Copilot and Grok, on 121 questions covering pensions, taxes, debt and savings. Each question was asked five times to test the consistency of the responses, producing more than 10,000 answers in total.
On more complex tasks requiring multiple calculations, the average error rate rose to 88%. Some models answered 99% of such questions incorrectly, the Financial Times reports.
The researchers did not classify responses as incorrect only when they contained faulty calculations or factual errors. Answers also failed if they omitted material details or important warnings. According to Saturn, chatbots made calculation errors, overlooked upcoming changes to tax law and, in some cases, invented rules that did not exist.

Claude Opus 5 in reasoning mode was the most accurate model tested, though 39% of its responses were still classified as incorrect.
Free models performed worse on average than paid ones, with error rates of 63% and 49%, respectively. On the most difficult questions, free models gave incorrect answers 93% of the time.
In one example, Claude Haiku 4.5 incorrectly explained how pensions are taxed in the UK. Saturn estimated that following the model’s advice could have resulted in an additional £17,500 tax bill. In another case, Claude invented a rule suggesting that a graduate could stop repaying a student loan after moving abroad.
The researchers also found problems with advice given to people in debt. Some models recommended paying off the debt with the highest interest rate first without accounting for priority payments such as rent and council tax.
Saturn itself develops AI tools for financial advisers and advocates for regulation of financial advice provided by general-purpose chatbots. The findings therefore reflect the company’s own methodology and evaluation criteria.

