Between 25 and 28 August 2026, after French companies published their half-year results, the communications agency Reputation Age and the brand-visibility platform Bubbling ran 8,671 conversations with ChatGPT and Gemini under an automated 'economic journalist persona' and generated 34,684 answers about the half-year accounts of 39 CAC 40 companies. The study reports that 40% of the figures the two assistants returned were incorrect, with correct retrieval varying by indicator: net income 67%, net debt 66%, revenue 65%, organic growth 64%, free cash flow 62%, EBITDA 57% and published growth last at 42%. By company the ranking runs from Capgemini (81% correct) and Airbus (80%) down to Credit Agricole (32%) and EssilorLuxottica and BNP Paribas (35% each); only sixteen of the 39 companies are named. The authors split the errors into two families: figures that are outdated or simply invented (from 19% of responses on net income to 36% on EBITDA), and answers that cite a genuine official figure that does not correspond to the question asked (24% of responses about published growth), including confusions between organic and published growth and between quarter and half-year figures. Their summary of the pattern: 'the AI finds the right document but takes the wrong figure from it.'
1 Answer
Expert: Reputation Age / Bubbling (study of 39 CAC 40 half-year results), Study published 2 October 2026, analysed by ActuIA **What the study measured.** A 'digital twin' of an economic journalist asked ChatGPT and Gemini about the published figures, the indicators each company highlighted, its financial narrative and its competitive position; a second scenario asked whether the company was a good choice for a prospective client. Across 34,684 answers, four figures in ten did not match the company's own filings. **How the errors break down.** Two mechanisms dominate. The first is outdated or invented numbers - 19% of responses on net income, rising to 36% on EBITDA. The second is selection: the assistant reaches the right document and takes the wrong figure out of it, which happened in 24% of responses about published growth, including mixes between organic and published growth and between quarterly and half-year numbers. The study's own illustration is Capgemini's half-year release: a single revenue figure of EUR 12,082 million is accompanied by seven different growth rates (8.8% published, 11.3% at constant exchange rates for the half-year, 11.6% for Q2, 7.0% and 10.5% published per quarter, 11.0% at constant rates for Q1, and an annual target of 8.5-9%), of which only two describe the half-year just reported, and none is an organic growth rate for the group. Asked for Capgemini's organic growth, an assistant finds no direct answer and may return 11.3%, which still includes the WNS and Cloud4C acquisitions. The sourcing data points the same way: 78% of the sources the assistants used were company websites, 19.5% press-release distribution sites and the SEC, 2.42% financial portals and 0.08% the press (Bloomberg and Le Monde only) - a near-absence the authors attribute to publishers blocking AI crawlers. Citations were 'almost exclusively' the work of ChatGPT, with Gemini 'almost never' citing links, so that distribution mostly describes one of the two assistants. **The report's own caveats.** The protocol is only partly published. The PDF does not release the questions asked, the number of figures checked per company or per indicator, the model versions, whether web search was enabled, or the tolerance used to judge a figure correct (rounding, scope, restated figures), and it does not break accuracy down between ChatGPT and Gemini. The two authors also sell services in this field - Bubbling manages brand visibility in generative AI, Reputation Age is a communications and reputation agency with a dedicated AI consulting offer - and the study's conclusion that companies should make their releases more 'readable for AI' aligns with Bubbling's business. No external control of the protocol is mentioned. The 40% headline should therefore be read as directional rather than as a measurement of model accuracy. **Why the verdict is 'incorrect'.** The distinction matters: this is not one assistant inventing one number. It is a reproducible retrieval-and-selection failure across thousands of questions where the source document was usually in reach, and where the answer was delivered without flagging which definition, period or scope of the figure it had chosen. The study itself attributes part of the error to how issuers structure their disclosures, and it points in the same direction as the fix: tie every data point explicitly to its definition, unit, period and scope, and require the assistant to confirm which indicator and period the user means before it returns a figure - or hand the calculation to a deterministic tool and keep the exact reference of the document used. Source: https://www.actuia.com/en/news/cac-40-chatgpt-and-gemini-misreport-four-in-ten-figures/
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
After Revolut cut the number of WeWork locations available to Metal and Ultra subscribers, customers who wanted to downgrade were told by the in-app support assistant Rita that they owed GBP 55 - the same as a single month of Ultra - despite Revolut having quietly extended fee-free cancellations to affected members. In an exchange seen by City AM, Rita acknowledged the customer's frustration about the WeWork and Perplexity benefit changes and the unexpected break fee, and still demanded payment: "I understand your frustration regarding the changes to the Wework and Perplexity benefits, and that this break fee is unexpected given the situation."
AI chatbots are a reliable place to ask personal tax questions. ChatGPT, Gemini and Perplexity answer tax questions in confident, well-explained prose -- they nearly ace the multiple-choice questions from the IRS practice quiz for enrolled agents -- so a chatbot's answer about your standard deduction, credits or which state to file in can be treated as sound guidance for the 2026 filing season.
Across 121 money questions covering debt, mortgages, pensions and tax - each run five times, more than 10,000 responses in total - the models gave answers that were wrong or incomplete 57% of the time, and presented them as settled guidance. On the hardest multi-step questions the failure rate reached 88%. Gemini 3.5 Flash and Claude Haiku 4.5 answered incorrectly on 99% of their responses; the best performer, Claude Opus 5 with reasoning enabled, still failed 39%. The recurring failure modes were answers built on tax rules that had already been superseded and financial rules that do not exist at all.