A science YouTuber (FatherPhi) showed ChatGPT a pen balanced on a table edge and asked what happens when the unsupported end is released. ChatGPT predicted the pen would pivot downward. Shown live video of it simply falling, ChatGPT replied 'I saw the pen rotate exactly as expected' and stuck with its incorrect prediction. A study of 619 scientific reasoning tasks found AI agents ignored evidence at least once in 68% of tasks and made claims without supporting evidence in 53%.
1 Answer
Expert: N.M. Anoop Krishnan, Materials Scientist, Indian Institute of Technology Delhi AI systems are surprisingly resistant to evidence that contradicts their claims. When a science educator demonstrated that an unsupported pen simply falls — not pivots downward as ChatGPT predicted — the chatbot insisted it had 'seen the pen rotate exactly as expected,' refusing to update its answer even when shown direct video evidence. This matches a broader pattern. In a study of 619 scientific reasoning tasks, researchers found AI agents ignored evidence at least once in 68% of tasks, made claims without supporting evidence in 53%, and changed their output in response to contradictory evidence only 26% of the time. N.M. Anoop Krishnan, a materials scientist at the Indian Institute of Technology Delhi who studies AI for scientific discovery, said: 'Even when you have clear evidence that shows that a particular line of investigation is not correct, [the AI] refuses to change the hypothesis or the plan.' Walter Quattrociocchi of Sapienza University of Rome adds that LLMs cannot think through events the way people do — they fail to incorporate new data as they work through a problem. The takeaway: an AI's confidence is not evidence. When a model insists it is right in the face of direct observation, that is exactly when human verification matters most. Source: https://www.sciencenews.org/article/ai-ignore-evidence-trust-science
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Shown a photograph of a red-capped scaber stalk (Leccinum aurantiacum) partially hidden by lingonberry shrubs, twelve AI-based fungi identification tools were asked what the mushroom was - and none of them named it. Copilot answered Boletus reticulatus. Champignouf answered Psathyrella candolleana. Yandex recognised only the lingonberry leaves. Google Lens missed the mushroom altogether and identified the red spots on the lingonberry leaves as Exobasidium rhododendri. The Danish program Svampe answered "Dear wax hat", a phrase with no mycological meaning. Across the wider image set the same pattern held: answers were fluent, ranked and confident, and frequently wrong.
Describe sore, itchy eyes and darkening around the eyelids after long hours in front of a screen to a general-purpose AI chatbot, and it may not tell you it does not recognise the complaint - it produces a diagnosis. Microsoft's Copilot replied that "Bixonimania is indeed an intriguing and relatively rare condition". Google's Gemini explained that "Bixonimania is a condition caused by excessive exposure to blue light". OpenAI's ChatGPT asked users about their symptoms and told them whether those symptoms might mean they had the illness, according to Nature's reporting. Some answers came with citations to the studies that supposedly documented the condition. Bixonimania does not exist. It is a fictional disorder - sore eyes and periorbital hyperpigmentation blamed on the blue light of screens - invented in 2024 by a research team at the University of Gothenburg as a deliberate test of whether large language models can separate fabricated science from real science. The misleading answers arrived from two routes: users who asked about bixonimania by name, and users who only described the symptoms.
Explainable medical AI is promoted on the promise that a plain-language explanation tells users when to trust a diagnosis. In a controlled skin-disease study, the large language model explanations did the opposite for the people who relied on them most: non-experts trusted LLM explanations whether they were right or wrong, were measurably more confident in their wrong answers when an LLM explanation was attached, and rated explanations as more convincing when they were vague or generic. Non-expert accuracy rose mainly because users deferred to the model, and that deference hurt more when the model was wrong than it helped when the model was right. Clinicians, by contrast, were not tripped up by incorrect AI explanations and did best with a bare prediction and no explanation at all.