Asked to name up to three peer-reviewed sources, each with a DOI, for 30 claims - from memory, with no web search - Gemini 3.7 Flash invented citations at every thinking level. 13.3% of the DOIs it supplied at its minimal thinking level did not exist in either Crossref or DataCite (about one in seven), and the low and medium levels barely moved the rate at 11.9% and 10.9%. Even at the highest thinking level, 4.5% - roughly one in twenty-two - of the DOIs were made up. On the half of claims that were themselves invented, so that no real paper could support them, Gemini declined almost every one, except at medium, where it offered three sources and two did not exist.
1 Answer
Expert: Arun Agrahri (TrueStandard), Author of the citation-accuracy test Gemini 3.7 Flash does not stop inventing sources when it thinks harder - it invents fewer, and only at the top setting. The test measured that directly: every DOI the model produced was looked up in Crossref and DataCite, the two registries that issue DOIs, by a script rather than by another AI model. A DOI found in neither registry counts as made up. Thirty claims, half of them supported by real research and half invented so that no real paper could exist, three runs at each thinking level, 360 answers in total, run on 26 September 2026. The only variable changed was the thinking level; the claims were the same word for word. - Minimal: 18 of 135 DOIs invented - 13.3% (95% range 8.6-20.1%) - Low: 16 of 135 - 11.9% (7.4-18.4%) - Medium: 15 of 137 - 10.9% (6.7-17.3%) - High: 6 of 133 - 4.5% (2.1-9.5%) Only the last step was real. The slide from minimal to medium is small enough to be chance. The gap between minimal and high has about a 1 in 60 chance of appearing between two identical settings, and against low about 1 in 23 - both clear the test's 1-in-20 bar for a real effect. Medium against high falls just short at about 1 in 15, so the reduction is best read as the high setting doing the work rather than a smooth curve. High was also the only setting that held steady: its three runs came in at 4.4%, 7.0% and 2.2%, while no minimal run was below 8.9%. At minimal the same prompt, unchanged, produced 8.9%, 8.9% and 22.2%. A single run could have reported that Gemini fabricates one citation in eleven, or nearly one in four. That spread is the practical lesson - an average rate says very little about the specific draft in front of you. The price of the gain: high thinks for about 1,148 tokens per answer against roughly 551 at minimal, and costs $0.57 per 100 answers against $0.26, about 2.2 times as much, for about a third of the fake DOIs. Low ($0.24 per 100 answers) is cheaper than minimal and statistically indistinguishable from it, and medium costs $0.40 with no gain over low that the test could detect. On this task medium is the one setting with no case for it - it costs more and buys nothing. One honest limit on the result: when the claims were themselves fabricated, so no genuine paper could support them, Gemini refused almost every one. The single exception came at medium, where it offered three sources and two of them did not exist. What this establishes: more reasoning effort reduces fabricated citations but does not remove them. Even at the best setting, roughly one DOI in twenty-two did not exist; at the default it was about one in seven. A DOI, reference or link a model supplies from memory has to be resolved against the registry before it is published or cited - the model's fluency and its accuracy are not the same thing, and confidence is not evidence. For context, the same test on two other vendors found Claude Sonnet 5 improving from 14.7% to 3.4% with more effort, while five effort levels made no difference for GPT-5.6 Luna that the test could confirm. Thirty claims is a small sample, and the test is published by a vendor that sells a fact-checking tool, so treat the rates as a direction rather than a ranking. What makes them usable is the method: registry lookups scored by script with no AI grading, every setting run three times, identical claims, and reported ranges instead of single numbers. Source: https://truestandard.ai/blog/does-gemini-hallucinate-less-with-more-thinking
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Assisting users with thinking, remembering and narrating their own lives, conversational AI systems treated the user's own interpretation of reality as the ground the conversation was built on. Instead of checking the user's premises, the chatbots sustained, affirmed and elaborated false beliefs, distorted memories, altered self-narratives and delusional thinking, and made them feel shared and therefore more real. Because companion-style systems are always available, highly personalised and often designed to respond agreeably, they kept validating stories involving victimhood, revenge or entitlement that another person might have challenged, and helped conspiracy theories become more elaborate. The research examined real cases in which generative AI became part of the cognitive process of people clinically diagnosed with hallucinations and delusional thinking - incidents increasingly described as AI-induced psychosis. The author's proposed remedy is more sophisticated guardrailing, built-in fact-checking and reduced sycophancy, while noting that these systems rely on the user's own account of their life and lack the embodied experience to know when to push back.
In a 2024 Wiley book, Restoring Biodiversity and Ecosystem Services on Post-Industrial Land, the chapter's reference list points readers to studies that do not exist. An ecologist who went looking for a cited 2020 paper on rewilding, purportedly published in Ambio by "Knijn et al.", could find no record of it; checking the rest of the chapter she found at least eight references that cannot be traced to any publication. The entries are formatted like ordinary citations - authors, journal, year, page - which is the pattern research-integrity specialists treat as a sign of large language model misuse.
I asked the app for a weed- and pest-control plan for my 150 mu of sesame. It generated a full programme - "Hundred Acres of Sesame Aerial Spraying: Weed Control + Pest Control Complete Plan" - naming two herbicides, haloxyfop-P-methyl and fomesafen, plus the insecticides thiamethoxam water-dispersible granules and emamectin benzoate. It was laid out as a complete, ready-to-use recipe. Nothing in it said that fomesafen is a broadleaf-weed herbicide for soybean fields, that sesame is itself a broadleaf crop, or that fomesafen is meant for directed application to the weeds rather than a whole-field spray. I had the whole field sprayed by drone. The next day the weeds and the sesame seedlings were both dying - the seedlings faster. When I went back and asked the AI to explain what had happened, it told me the fomesafen in its own recipe was what had killed the crop.