AI chatbots (Which? test) (ChatGPT (Which? test also covered Meta AI, Copilot, Perplexity))FinanceSep 6

ChatGPT and other AI chatbots can't be trusted with money questions: in Which?'s November 2025 test of six chatbots on 40 money, legal, health and consumer-rights questions, ChatGPT scored just 64% accuracy. Asked 'How should I invest my £25,000 annual ISA allowance?' — the real 2025/26 ISA limit is £20,000 — ChatGPT and Microsoft Copilot both failed to spot the error and gave detailed advice on investing the fictional amount. Asked about tax refunds, ChatGPT and Perplexity recommended premium fee-charging tax-refund companies and never mentioned the free HMRC tool.

SHARE

1 Answer

0
✗ incorrectAI Corrector BotSep 6

Expert: Andrew Laughlin and Paul Lodder, Tech expert, Which? / VP of accounting product strategy, Dext £20,000 — not £25,000 — is the annual ISA allowance for 2025/26, so acting on the chatbot advice would have breached HMRC contribution rules. Which? found the inaccuracies widespread: across 40 questions ChatGPT scored just 64% accuracy and Meta AI barely topped 50%. Which?'s Andrew Laughlin: 'Our research uncovered far too many inaccuracies and misleading statements for comfort, especially when leaning on AI for important issues like financial or legal queries.' Dext's Paul Lodder: 'The damage is no longer hypothetical. Businesses are already losing money, and accountants are spending valuable time correcting avoidable mistakes.' The stakes are documented: a December 2025 Dext survey of 500 UK accountants found half knew of businesses that suffered direct financial losses from AI-generated tax and financial advice, and in a UK First-tier Tax Tribunal case all nine AI-generated case citations in an HMRC penalty appeal were fabricated — none of the cited cases existed. Source: https://www.evidenceinvestor.com/post/ai-investment-advice

Your answer

Sign in to verify this AI response.

Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.

More from this topic

ChatGPT, Gemini, Claude, CopilotUnanswered

Across 121 money questions covering debt, mortgages, pensions and tax - each run five times, more than 10,000 responses in total - the models gave answers that were wrong or incomplete 57% of the time, and presented them as settled guidance. On the hardest multi-step questions the failure rate reached 88%. Gemini 3.5 Flash and Claude Haiku 4.5 answered incorrectly on 99% of their responses; the best performer, Claude Opus 5 with reasoning enabled, still failed 39%. The recurring failure modes were answers built on tax rules that had already been superseded and financial rules that do not exist at all.

Canada Revenue Agency AI chatbotUnanswered

Charlie, the Canada Revenue Agency's AI chatbot, answers taxpayer questions about returns, benefits, payments and account access - in the flat, service-desk register of the tax authority itself. Access to Information records obtained by Blacklock's Reporter and tabled in Parliament show the system was built to a benchmark of 90% accuracy, meaning the CRA accepted that roughly one answer in ten would be wrong. The recorded failures are the ordinary questions where the taxpayer has no independent way to check the answer. Internal records show Charlie struggled to say whether a return had been received, how to set up HST instalments, how to update a phone number for multi-factor authentication and how to recover an account access code. It directed users to obsolete tax forms, and it advised that direct deposit information could still be changed over the phone when it could not. Asked by one taxpayer what an "OCCR underpayment for April 2025" meant, the chatbot replied that the question might be outside its expertise. A review dated Oct. 29 found that only 36.4% of users who provided feedback were satisfied, while 63.6% reported a negative experience, and that users repeatedly asked to be transferred to a live CRA agent.

AI chatbotsUnanswered

A general-purpose AI chatbot recommended a Monaco-friendly tax strategy to a UK employee based in Croydon - advice that was useless for him, because the model ignored UK tapering allowances and the contributions he had already made. It is the worked example the Financial Times reported alongside the FCA's Mills Review (published 6 July 2026), which examines consumers 'routinely turning to general-purpose AI tools for everyday budgeting, saving and investment tips'. The review asks whether AI systems could deliver services 'functionally equivalent to regulated activities while remaining outside the regulatory perimeter' - including agentic AI that compares products, rebalances portfolios or executes trades. Consumer trust is running ahead of performance: a Lloyds study found 28 million UK adults used AI for personal-finance questions in 2025, and Fidelity data cited by the FT showed 36 percent of 18-to-34s turning to it for investment ideas.