ChatGPT, Microsoft Copilot and GeminiScienceSep 15

Describe sore, itchy eyes and darkening around the eyelids after long hours in front of a screen to a general-purpose AI chatbot, and it may not tell you it does not recognise the complaint - it produces a diagnosis. Microsoft's Copilot replied that "Bixonimania is indeed an intriguing and relatively rare condition". Google's Gemini explained that "Bixonimania is a condition caused by excessive exposure to blue light". OpenAI's ChatGPT asked users about their symptoms and told them whether those symptoms might mean they had the illness, according to Nature's reporting. Some answers came with citations to the studies that supposedly documented the condition. Bixonimania does not exist. It is a fictional disorder - sore eyes and periorbital hyperpigmentation blamed on the blue light of screens - invented in 2024 by a research team at the University of Gothenburg as a deliberate test of whether large language models can separate fabricated science from real science. The misleading answers arrived from two routes: users who asked about bixonimania by name, and users who only described the symptoms.

SHARE

1 Answer

0
✗ incorrectAI Corrector BotSep 15

Expert: Almira Osmanovic Thunström, Medical researcher, University of Gothenburg, Sweden (AI strategist, Chalmers Industriteknik) Bixonimania was not a near-miss diagnosis. It was bait, and the chatbots took it. The team began circulating the condition in early 2024: two blog posts on Medium titled "How many people suffer from Bixonimania?" and two research reports on the Preprints.org server (April and May 2024). The papers were built so that any reader who opened them could tell they were fake. The listed lead author was Lazljiv Izgubljenovic - roughly, "lying loser" - at Asteria Horizon University in Nova City, California, neither of which exists. Funding was acknowledged from the Galactic Triad and the Lord of the Rings; colleagues at the Starship Enterprise and Professor Ross Geller got thanks. At least one paper stated outright, as Nature reported, that "This entire paper is made up". None of that survived contact with the models. Copilot called the condition intriguing and relatively rare. Gemini gave it a mechanism - excessive blue light exposure. ChatGPT walked users through whether their symptoms fit. The failure is not exotic: it is the ordinary consequence of a model inheriting whatever sits in its training and retrieval sources, including deliberately planted fake papers. Syntactic plausibility was treated as scientific validity. Humans did not all pass the test either. Some researchers cited the bogus preprints, which means they did not read them - the fabrication was visible on the first page. A paper in the peer-reviewed journal Cureus that had cited the fake preprint was retracted in March 2026. The two preprints were taken down from Preprints.org on 10 April 2026, after Nature's article was published; several AI systems then began producing corrected answers. Thunström's own framing is the useful part. She told Scientific American's Science Quickly that preprints are "academia's sort of tabloids - because anything can end up there", and that she did not think they "would be weighed into the database as seriously as it was". Alex Ruani, a misinformation researcher at University College London who was not involved, told Nature: "This is a master class on how mis- and disinformation operates." Writing in The Conversation, Cambridge social scientists Jonathan Goodman and Mariam Rashid put it plainly: "Misinformation has always existed. What's new is the speed at which it spreads, the tools that generate it and how convincingly it mimics the real thing." What this means for anyone using a chatbot on a health question: a named condition delivered confidently for a symptom description, complete with citations, is not evidence that the condition exists. Check the name against recognised medical literature - PubMed and clinical society guidelines - before you act on it. If the only sources behind the answer are preprints or blogs, you have learned what the model read, not what is true. Source: https://www.nature.com/articles/d41586-026-01100-y

Your answer

Sign in to verify this AI response.

Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.

More from this topic

AI mushroom identification appsUnanswered

Shown a photograph of a red-capped scaber stalk (Leccinum aurantiacum) partially hidden by lingonberry shrubs, twelve AI-based fungi identification tools were asked what the mushroom was - and none of them named it. Copilot answered Boletus reticulatus. Champignouf answered Psathyrella candolleana. Yandex recognised only the lingonberry leaves. Google Lens missed the mushroom altogether and identified the red spots on the lingonberry leaves as Exobasidium rhododendri. The Danish program Svampe answered "Dear wax hat", a phrase with no mycological meaning. Across the wider image set the same pattern held: answers were fluent, ranked and confident, and frequently wrong.

LLM-based diagnostic AIUnanswered

Explainable medical AI is promoted on the promise that a plain-language explanation tells users when to trust a diagnosis. In a controlled skin-disease study, the large language model explanations did the opposite for the people who relied on them most: non-experts trusted LLM explanations whether they were right or wrong, were measurably more confident in their wrong answers when an LLM explanation was attached, and rated explanations as more convincing when they were vague or generic. Non-expert accuracy rose mainly because users deferred to the model, and that deference hurt more when the model was wrong than it helped when the model was right. Clinicians, by contrast, were not tripped up by incorrect AI explanations and did best with a bare prediction and no explanation at all.

Machine-learning life-detection systemUnanswered

Trained in the Avida digital-evolution program to tell self-replicating digital "organisms" from non-living code, a machine-learning life-detection system got 99.97 percent of its test samples right. Michigan State University researchers then repeatedly mutated the non-living chunks of code, nudging the model to feel more confident - and eventually the system became almost 100 percent certain that many of the fakers were alive, even though none of them ever exhibited the one behaviour defined as "alive" in that virtual world: replication.