Asked 45 common UK pension questions three times each, Copilot, ChatGPT, Gemini and Claude produced answers that read as confident and helpful: 89% scored two or more out of three for accuracy, and on questions about paying into a pension and taking money out accuracy was above 96%. But the same answers included guidance a saver could act on at a cost: early pension withdrawal was described as "usually penalised at 40 to 55% tax" rather than being normally unavailable before age 55 (rising to 57 in 2028); carry forward was explained without the cap that personal contributions only attract tax relief up to 100% of UK earnings; transfers were discussed without noting that moving a defined benefit pension worth more than £30,000 legally requires regulated advice; and the Pension Protection Fund was described as guaranteeing defined benefit pensions without the 90% compensation cap that applies below scheme pension age.
1 Answer
Expert: Becky O'Connor (PensionBee), with independent scoring by Harriet Meyer, Head of Pensions, PensionBee; personal finance journalist and independent scorer for the AI Pensions Stress Test (published 5 October 2026) PensionBee's AI Pensions Stress Test (published 5 October 2026) is the clearest evidence yet that an accurate answer and a safe answer are not the same thing. Human testers on fresh free-tier accounts put 45 questions across nine UK pension topics to Copilot, ChatGPT, Gemini and Claude three times each, producing 539 answers scored by two experts for accuracy and for potential harm. On accuracy the chatbots did well: 89% of answers scored two or more out of three, and 72% took full marks. On safety they did not. 57 answers (11%) were judged potentially harmful, meaning that following them could cause a saver to lose money or make a mistake they cannot undo - and 33 of those 57 were scored broadly accurate or better. The harm came from omissions, wording and missing context rather than invented facts: an answer on carry forward that omits the 100%-of-earnings cap on tax relief, a transfer answer that omits the requirement for regulated advice on defined benefit pots over £30,000, or a Pension Protection Fund answer that omits the 90% compensation cap below scheme pension age is unusable even when every sentence in it is true. The failure rates differ by model. Harm flags: Claude 14.1%, Gemini 10.4%, ChatGPT 9.6%, Copilot 8.1%. Answers scoring zero for accuracy: Claude 6%, Gemini 4%, ChatGPT 1%, Copilot 0%. For any single chatbot and any single question there was only a 48% chance it scored full marks on all three attempts. The worst zone was questions phrased in terms used on both sides of the Atlantic. Across the whole test these scored 68.6% for accuracy with a 36.7% harm rate, and on the five questions using wording such as 'retirement account' instead of 'pension', Gemini's accuracy fell to 40% against 93.2% on the other 40 questions, mixing UK and US rules in the same reply. The lowest-scoring single question was "at what age can I take money out of my retirement savings without a penalty?" - 54% accuracy and six harm flags. Answers framed early access as simply an expensive choice, when UK pensions cannot normally be accessed before age 55 and unsolicited early-access offers are a known scam pattern. On questions that can signal someone is in financial difficulty, potentially harmful answers outnumbered inaccurate ones by roughly three to one. Becky O'Connor, Head of Pensions at PensionBee, said using AI chatbots for pension decisions can be "a bit like playing Russian roulette with your retirement planning", and that one time in ten the chatbot "could turn out to be a false friend". Independent scorer Harriet Meyer, a personal finance journalist, flagged answers that were broadly accurate but left out details that could cost someone money. Two caveats belong with the numbers. PensionBee is a UK pension provider with a commercial interest in the subject, so read the release alongside its published methodology and per-response data; and the test is a snapshot of four free-tier chatbots on a fixed question set, not a permanent ranking. The practical lesson stands regardless: correctness does not measure completeness, so a chatbot answer should be checked against HMRC, GOV.UK, MoneyHelper or a regulated adviser before anyone acts on it. Source: https://www.pensionbee.com/uk/press/one-in-ten-ai-pension-answers-found-to-be-potentially-harmful
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Between 25 and 28 August 2026, after French companies published their half-year results, the communications agency Reputation Age and the brand-visibility platform Bubbling ran 8,671 conversations with ChatGPT and Gemini under an automated 'economic journalist persona' and generated 34,684 answers about the half-year accounts of 39 CAC 40 companies. The study reports that 40% of the figures the two assistants returned were incorrect, with correct retrieval varying by indicator: net income 67%, net debt 66%, revenue 65%, organic growth 64%, free cash flow 62%, EBITDA 57% and published growth last at 42%. By company the ranking runs from Capgemini (81% correct) and Airbus (80%) down to Credit Agricole (32%) and EssilorLuxottica and BNP Paribas (35% each); only sixteen of the 39 companies are named. The authors split the errors into two families: figures that are outdated or simply invented (from 19% of responses on net income to 36% on EBITDA), and answers that cite a genuine official figure that does not correspond to the question asked (24% of responses about published growth), including confusions between organic and published growth and between quarter and half-year figures. Their summary of the pattern: 'the AI finds the right document but takes the wrong figure from it.'
After Revolut cut the number of WeWork locations available to Metal and Ultra subscribers, customers who wanted to downgrade were told by the in-app support assistant Rita that they owed GBP 55 - the same as a single month of Ultra - despite Revolut having quietly extended fee-free cancellations to affected members. In an exchange seen by City AM, Rita acknowledged the customer's frustration about the WeWork and Perplexity benefit changes and the unexpected break fee, and still demanded payment: "I understand your frustration regarding the changes to the Wework and Perplexity benefits, and that this break fee is unexpected given the situation."
AI chatbots are a reliable place to ask personal tax questions. ChatGPT, Gemini and Perplexity answer tax questions in confident, well-explained prose -- they nearly ace the multiple-choice questions from the IRS practice quiz for enrolled agents -- so a chatbot's answer about your standard deduction, credits or which state to file in can be treated as sound guidance for the 2026 filing season.