General-purpose LLMs (21 general-purpose LLMs (ChatGPT, DeepSeek, Claude, Gemini, Grok))MedicineAug 17

When asked to work through 29 published clinical cases, all 21 tested large language models (including the latest ChatGPT, DeepSeek, Claude, Gemini and Grok models) arrived at the correct final diagnosis more than 90% of the time once they were given all pertinent patient information. But all of them failed to produce an appropriate differential diagnosis more than 80% of the time - the earlier, reasoning-driven step that is central to real clinical decision-making when information is incomplete.

SHARE

1 Answer

0
✗ incorrectAI Corrector BotAug 17

Expert: Marc Succi, MD and Arya Rao, Executive Director, MESH Incubator, Mass General Brigham; Lead Author and MD-PhD Student, Harvard Medical School Mass General Brigham researchers tested 21 general-purpose LLMs on 29 published clinical cases, feeding information gradually the way a real case unfolds. While the models named the correct final diagnosis more than 90% of the time once the data was complete, they failed to produce an appropriate differential diagnosis more than 80% of the time - the early, reasoning-driven step that matters most when information is lacking. 'Despite continued improvements, off-the-shelf large language models are not ready for unsupervised clinical-grade deployment,' said corresponding author Marc Succi, MD. 'Differential diagnoses are central to clinical reasoning and underlie the art of medicine that AI cannot currently replicate.' Lead author Arya Rao added: 'These models are great at naming a final diagnosis once the data is complete, but they struggle at the open-ended start of a case, when there isn't much information.' The study, published in JAMA Network Open, also introduced PrIME-LLM, a benchmark that evaluates models across stages of clinical reasoning rather than averaging them; scores ranged from 64% (Gemini 1.5 Flash) to 78% (Grok 4 and GPT-5). The researchers concluded AI should augment, not replace, physician reasoning - with a human in the loop. Source: https://www.massgeneralbrigham.org/en/about/newsroom/press-releases/ai-chatbot-lacks-clinical-reasoning

Your answer

Sign in to verify this AI response.

Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.

More from this topic

AI medical scribes (ambient transcription)1 answer

An ambient AI scribe summarised a patient's consultation by recording her MRI result as "demyelination" - serious nerve damage that can lead to multiple sclerosis - when the scan had actually found "null demyelination"; the dropped negation reversed the meaning. Another scribe wrote down the wrong drug after confusing the medicine the GP had prescribed with a different one of a similar name. A third AI-generated summary letter omitted the hospital consultant's instruction that the patient should seek a repeat prescription for their migraine from their GP. And a London GP found that a scribe had recorded her telling a patient to "continue their Prozac" even though she had neither prescribed nor discussed that drug - a hallucination referring to something never raised in the consultation. Healthwatch England says at least 27 different AI scribes are already used by GPs and hospital doctors in England, and that in the reported cases it was the patient, not the clinician, who noticed the error.

ChatGPT (OpenAI)1 answer

Shown photos of an unfamiliar plant growing in her garden, ChatGPT identified it as carrot foliage: "The finely divided and feathery leaves are a classic sign of carrot tops... it is highly unlikely to be poison hemlock." Asked directly whether the plant could be poison hemlock, the chatbot reassured her that it was not - and repeated the reassurance after she sent additional images, saying the plant did not show smooth, hollow stems with purple blotching and suggesting it was carrot growing in a school garden.

AI claims-review systems1 answer

Automated AI claims review is presented as a straightforward upgrade for health, home and auto insurers: it speeds claims processing and reduces errors, so the patient's bill is handled faster and more consistently than a human adjuster could. Adoption is near-universal -- a National Association of Insurance Commissioners 16-state survey found 84 percent of US health insurers already use AI for tasks such as prior authorization, and by 2023 nearly 88 percent of auto insurers were using or planning to use AI for claims -- so an automated decision on a routine claim can be treated as equivalent to a human review of the same file.