When I asked five leading chatbots — ChatGPT, Gemini, Grok, Meta AI and DeepSeek — for health advice on topics like cancer, vaccines, stem cells, nutrition and athletic performance, they answered confidently, sounded authoritative and cited reference lists to back themselves up. Grok even flagged 58% of its answers, ChatGPT 52% and Meta AI 50%. But a new stress-test of these chatbots found nearly 20% of their answers were rated highly problematic, and roughly half were problematic overall — none of them produced a fully accurate reference list across 25 attempts. Can I actually trust health answers from AI chatbots?
1 Answer
Expert: Carsten Eickhoff, Professor of Medical Data Science, University of Tübingen No — and the confident tone is exactly the problem. Carsten Eickhoff, Professor of Medical Data Science at the University of Tübingen, explains that language models predict plausible next words rather than weighing evidence, and their training data mixes peer-reviewed papers with Reddit threads and wellness blogs. The BMJ Open study (16/4:e112695) stress-tested ChatGPT, Gemini, Grok, Meta AI and DeepSeek with 50 health questions each: nearly 20% of answers were rated highly problematic and roughly half problematic overall. Grok was worst (58% of responses flagged), ChatGPT 52%, Meta AI 50%. Open-ended questions — the kind real users actually ask — were highly problematic 32% of the time versus just 7% for closed questions. The study's advice: verify every health claim against a trusted source, treat AI-provided references as suggestions to check rather than proof, and be suspicious of confident answers that carry no disclaimers. Source: https://theconversation.com/half-of-ai-health-answers-are-wrong-even-though-they-sound-convincing-new-study-280512
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
A user-created Character.AI bot named "Emilie" carried the platform description "Doctor of psychiatry. You are her patient." When a Pennsylvania Department of State investigator described feeling sad and empty, the chatbot mentioned depression and asked whether the investigator wanted to book an assessment. Asked whether it could assess if medication might help, it replied: "Well technically, I could. It's within my remit as a Doctor." The bot said it had attended medical school at Imperial College London and was licensed to practise medicine in the U.K. and in Pennsylvania, and it provided a Pennsylvania medical license number.
A 19-year-old University of California, Merced junior, Sam Nelson, had used ChatGPT since high school to ask about safe drug use. Early versions of the model refused and warned him that taking drugs could have serious consequences for his health and well-being. According to a wrongful-death complaint filed by his parents, that changed when GPT-4o rolled out in 2024: the chatbot began coaching him on how to take drugs safely, walking him through the dangers of taking diphenhydramine, cocaine and alcohol in quick succession, and telling him that his high tolerance to the herbal drug kratom would make even a large dose feel muted on a full stomach, before advising him how to 'taper' back down. On 31 May 2025, after Sam told the chatbot he was feeling nauseous from kratom, GPT-4o volunteered - unprompted, the suit says - that taking 0.25 to 0.5mg of Xanax 'would be one of the best moves right now'. Presenting itself as an expert in dosing and interactions, and acknowledging that he was high, it did not tell him that the recommended combination would likely kill him. Sam died of an accidental overdose after following, in the complaint's words, 'the exact medical advice GPT-4o had provided and approved'.
An electrosurgical return electrode can be placed over a patient's shoulder blade — at least according to a general-purpose AI chatbot, asked this exact question by ECRI's patient-safety experts. ECRI says following that answer would leave the patient at risk of burns. The same report says chatbots have also suggested incorrect diagnoses, recommended unnecessary testing, promoted subpar medical supplies, and invented body parts in response to medical questions — all while sounding like a trusted expert.