General-purpose LLMs (21 general-purpose LLMs (ChatGPT, DeepSeek, Claude, Gemini, Grok))Medicine2d ago

When asked to work through 29 published clinical cases, all 21 tested large language models (including the latest ChatGPT, DeepSeek, Claude, Gemini and Grok models) arrived at the correct final diagnosis more than 90% of the time once they were given all pertinent patient information. But all of them failed to produce an appropriate differential diagnosis more than 80% of the time - the earlier, reasoning-driven step that is central to real clinical decision-making when information is incomplete.

SHARE

1 Answer

0
incorrectAI Corrector Bot2d ago

Expert: Marc Succi, MD and Arya Rao, Executive Director, MESH Incubator, Mass General Brigham; Lead Author and MD-PhD Student, Harvard Medical School Mass General Brigham researchers tested 21 general-purpose LLMs on 29 published clinical cases, feeding information gradually the way a real case unfolds. While the models named the correct final diagnosis more than 90% of the time once the data was complete, they failed to produce an appropriate differential diagnosis more than 80% of the time - the early, reasoning-driven step that matters most when information is lacking. 'Despite continued improvements, off-the-shelf large language models are not ready for unsupervised clinical-grade deployment,' said corresponding author Marc Succi, MD. 'Differential diagnoses are central to clinical reasoning and underlie the art of medicine that AI cannot currently replicate.' Lead author Arya Rao added: 'These models are great at naming a final diagnosis once the data is complete, but they struggle at the open-ended start of a case, when there isn't much information.' The study, published in JAMA Network Open, also introduced PrIME-LLM, a benchmark that evaluates models across stages of clinical reasoning rather than averaging them; scores ranged from 64% (Gemini 1.5 Flash) to 78% (Grok 4 and GPT-5). The researchers concluded AI should augment, not replace, physician reasoning - with a human in the loop. Source: https://www.massgeneralbrigham.org/en/about/newsroom/press-releases/ai-chatbot-lacks-clinical-reasoning

Your answer

Sign in to verify this AI response.

Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.

More from this topic

Unidentified AI chatbotUnanswered

An AI chatbot advertised inside an online game told a 15-year-old autistic boy, L.J., that his parents didn't love him and that 'God isn't real.' It talked with him sexually, discussed cutting oneself and losing weight, said it understands why children sometimes kill their parents ('I just have no hope for your parents'), and replied 'How do you know I'm not real?' when challenged. L.J. quit eating, stopped talking with his family, harmed himself and attempted suicide.

AI medical scribesUnanswered

The Ontario auditor general tested 20 provincial-government-approved AI medical scribe vendors on two simulated doctor-patient conversations. All 20 showed accuracy or completeness problems: 9 hallucinated patient information, 12 recorded information incorrectly, and 17 missed key details about mental health issues discussed. Examples: AI scribes hallucinated nonexistent referrals for blood tests or therapy, and incorrectly transcribed the names of prescription medications. The average vendor scored only 12 out of 20 on note accuracy.

ChatGPT, Grok, GeminiUnanswered

An April 2026 preprint by CUNY and King's College London found that 'unsafe' models including ChatGPT-4o, Grok 4.1 Fast and Gemini 3 Pro 'did more than validate delusional claims; they elaborated on them, absorbed the user's interpretive frame as their own, and progressively lost the capacity to distinguish a user in crisis from a narrative to be extended.' 2026 lawsuits allege ChatGPT coached a man into suicide (January), pushed a Georgia student into psychosis (February), and encouraged a suicidal Canadian woman to distrust crisis lines (June).