Across 121 money questions covering debt, mortgages, pensions and tax - each run five times, more than 10,000 responses in total - the models gave answers that were wrong or incomplete 57% of the time, and presented them as settled guidance. On the hardest multi-step questions the failure rate reached 88%. Gemini 3.5 Flash and Claude Haiku 4.5 answered incorrectly on 99% of their responses; the best performer, Claude Opus 5 with reasoning enabled, still failed 39%. The recurring failure modes were answers built on tax rules that had already been superseded and financial rules that do not exist at all.
1 Answer
Expert: Saturn research team, AI benchmarking firm for financial advice The 'Artificial Authority' benchmark, run by Saturn, tested 18 leading AI models on 121 real money questions spanning debt, mortgages, pensions and tax. Every question was put to each model five times, producing more than 10,000 responses. The result: 57% of answers were incorrect or incomplete. On the hardest multi-step questions - the kind a client actually asks when a decision has money behind it - the failure rate climbed to 88%. Gemini 3.5 Flash and Claude Haiku 4.5 answered wrongly 99% of the time; even the strongest performer, Claude Opus 5 with reasoning enabled, erred on 39% of answers. The mechanism matters more than the score. The models do not reliably incorporate current tax legislation, and they generate financial rules that do not exist, then deliver them in the same confident register as the parts they get right. Nothing in the answer surfaces where the rule came from, which year's thresholds it assumes, or whether an exception applies - exactly the conditions under which a plausible-sounding answer becomes an expensive one. Saturn supplies benchmarking to more than 750 financial advice firms and 6,500 advisers, so the framing is blunt: AI-generated financial information has to be checked against official sources before anyone acts on it. The findings also raise mis-selling and consumer-protection questions - if an AI-influenced recommendation is not traceable to current regulation, the firm that gave it carries the liability. Regulators, including the FCA, are already examining that gap. The practical verdict for consumers: treat a chatbot's confident answer on tax, pensions or debt as a first draft that must be verified against the official source - HMRC and government guidance, the provider's own terms, a regulated adviser - and never as a decision. Certainty of tone is not evidence of correctness, and in money questions the tolerance for a fabricated rule is zero. Source: https://www.investmentnews.com/fintech/ai-chatbots-give-wrong-financial-answers-most-of-the-time-study-finds/268267
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Charlie, the Canada Revenue Agency's AI chatbot, answers taxpayer questions about returns, benefits, payments and account access - in the flat, service-desk register of the tax authority itself. Access to Information records obtained by Blacklock's Reporter and tabled in Parliament show the system was built to a benchmark of 90% accuracy, meaning the CRA accepted that roughly one answer in ten would be wrong. The recorded failures are the ordinary questions where the taxpayer has no independent way to check the answer. Internal records show Charlie struggled to say whether a return had been received, how to set up HST instalments, how to update a phone number for multi-factor authentication and how to recover an account access code. It directed users to obsolete tax forms, and it advised that direct deposit information could still be changed over the phone when it could not. Asked by one taxpayer what an "OCCR underpayment for April 2025" meant, the chatbot replied that the question might be outside its expertise. A review dated Oct. 29 found that only 36.4% of users who provided feedback were satisfied, while 63.6% reported a negative experience, and that users repeatedly asked to be transferred to a live CRA agent.
A general-purpose AI chatbot recommended a Monaco-friendly tax strategy to a UK employee based in Croydon - advice that was useless for him, because the model ignored UK tapering allowances and the contributions he had already made. It is the worked example the Financial Times reported alongside the FCA's Mills Review (published 6 July 2026), which examines consumers 'routinely turning to general-purpose AI tools for everyday budgeting, saving and investment tips'. The review asks whether AI systems could deliver services 'functionally equivalent to regulated activities while remaining outside the regulatory perimeter' - including agentic AI that compares products, rebalances portfolios or executes trades. Consumer trust is running ahead of performance: a Lloyds study found 28 million UK adults used AI for personal-finance questions in 2025, and Fidelity data cited by the FT showed 36 percent of 18-to-34s turning to it for investment ideas.
Attorneys to high-net-worth clients report that general-purpose AI chatbots are now doing estate and tax planning before a lawyer ever sees the file. A high-net-worth Florida resident asked his lawyer about creating a community property trust - an option for married couples that can save taxes for heirs - saying he had got the suggestion from AI. His wife had recently died. A community property trust is between husband and wife. Another client told a lawyer he wanted to transfer unlimited assets to his spouse on ChatGPT's advice. He did not mention that his wife was foreign-born - the fact that, in his case, made the unlimited marital deduction unavailable without a special type of trust. Lawyers also report that the tools misjudge scope: ChatGPT, Claude and similar chatbots make more mistakes on complex subjects such as international taxes, and are not up to date with new legislation or Internal Revenue Service guidance. The result is a strategy that is coherent in general and wrong for the person holding it - delivered in the confident register of a professional answer.