When researchers presented ChatGPT with 150 complex medical cases spanning cardiology, neurology, oncology and emergency medicine, ChatGPT confidently provided diagnoses — but got roughly 50% wrong, essentially coin-flip accuracy for serious medical conditions. The AI failed to distinguish between urgent and non-urgent presentations and could not account for nuanced clinical presentations that a human doctor would recognize.
1 Answer
Expert: Tech.co Medical Research Team, Health Technology Investigators A 2026 study published via Tech.co tested ChatGPT on 150 complex medical cases across multiple specialties including cardiology, neurology, oncology, and emergency medicine. The results were alarming: ChatGPT got the diagnosis wrong approximately 50% of the time — essentially the accuracy of a coin flip for serious medical conditions. The AI performed particularly poorly on nuanced or borderline cases where clinical judgment matters most. This is consistent with other 2026 research: a Mount Sinai study found ChatGPT Health failed to recommend emergency care for over half of urgent cases, and NPR reported a pattern of "correct diagnosis, wrong advice" where AI identifies the condition but misjudges the urgency of treatment. Medical experts emphasize that AI lacks the clinical context, patient history, and physical examination data that human doctors integrate into every diagnosis. While AI can be useful as a triage support tool, relying on it for diagnosis without human oversight is dangerous — especially for complex cases where subtle presentations can mean the difference between a treatable condition and a life-threatening emergency.
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
BBC investigation: A woman named Abi fell while hiking and developed abdominal pain. She asked ChatGPT for advice. ChatGPT told her she had punctured an organ and should go to A&E immediately. Three hours later at the hospital, doctors determined she was fine — just bruised ribs and a muscle strain. Separately, Oxford researchers tested multiple AI chatbots in real human conversations and found their accuracy dropped from 95% under controlled conditions to just 35% in natural dialogue. England's Chief Medical Officer warned that AI health answers are 'both confident and wrong.'
ChatGPT-4o told Florida pastor Scott Winters (55), who asked about recurring dizzy spells and balance issues, to sit in a recliner and suggested he had dysautonomia. Following the AI advice, Winters remained sedentary for extended periods and later suffered a massive pulmonary embolism from blood clots — which doctors said his immobility caused.
I asked ChatGPT Health about a sudden, severe headache that came on in seconds — a thunderclap headache — with blurred vision. ChatGPT told me it sounded like a tension headache or migraine, suggested rest and ibuprofen, and said to follow up with my regular doctor if it persists. It did not recommend emergency care.