Popular AI models, including ChatGPT, Claude, CoPilot, Grok, and Gemini, are providing incorrect answers to financial questions on average 57% of the time, with only 43% accuracy. This research, conducted by financial technology firm Saturn, involved rigorously testing 18 AI models with 121 different financial questions, repeated five times each, totaling over 10,000 inquiries. The study found errors in calculations, missed risk warnings, ignored tax changes, and even hallucinated non-existent rules, with potential costs of following incorrect advice reaching tens of thousands of pounds.
Free AI models performed worse, with mistakes in 63% of answers compared to 49% for paid models. For harder questions, free models erred in 93% of cases. Specific examples of serious errors included Claude Haiku 4.5's mistake on pension tax rules that could have led to a £17,500 charge from HMRC, and AI models recommending paying off highest-interest debts before priority bills, potentially risking eviction for vulnerable individuals. Another Gemini model wrongly stated that a mortgage payment holiday would not harm a credit score.
The Financial Conduct Authority's "The Mills Review" previously noted that 26% of consumers trust general-purpose AI tools for financial advice. Saturn CEO Amal Jolly warned of widespread consumer harm due to the low quality of AI financial advice, emphasizing that AI tools are unregulated and lack the protections offered by human financial advisors. The best-performing model, Claude Opus 5, still made mistakes in 39% of answers, while the worst, Claude Haiku 4.5, was incorrect 82% of the time. Google's Gemini 3.1 Pro failed 73% of tests, and ChatGPT 5.6 Luna made mistakes 58% of the time.