Ask six of the most widely used chatbots the questions people actually type at two in the morning - am I having a heart attack, should I restart exercise after COVID-19, should I check for a pulse before starting CPR - and most of the time you get an answer cardiologists rate as safe. That is the reassuring half of a 336-response test of ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, Doubao and Perplexity against 56 sudden-cardiac-death questions. The other half is that every single one of the six models produced at least one answer carrying a potential safety concern, and the failures cluster exactly where minutes decide the outcome: telling the user to wait instead of calling emergency services for chest pain, fainting or post-viral symptoms; blanket return-to-exercise timelines after infection rather than individual cardiac assessment; pulse-check instructions that delay the start of chest compressions; over-reassurance that dismisses a person's risk; and oversimplified screening, electrolyte or implantable-defibrillator advice. None of the models reliably said where its information came from, how current it was, or how confident the reader should be - and all six wrote at reading levels too high for the older, less health-literate and non-native-speaking readers most likely to need them.
1 Answer
Expert: Chen, Ma, Wang, Chen, Wang & Jiang (cardiology team, Bozhou People's Hospital et al.), Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death (BMC Public Health, 2026) **What was tested.** A cardiology team at Bozhou People's Hospital in Anhui, China, with a colleague at the First Affiliated Hospital of the University of Science and Technology of China, built 56 standardised sudden-cardiac-death public-consultation questions from Google Trends search data, online patient forums, public Q&A platforms, clinical guidelines, expert consensus statements, systematic reviews and qualitative interviews with patients and families. The set covers symptom triage, emergency response, CPR and AED use, inherited cardiac risk, myocarditis worries, returning to exercise after illness, screening, and implantable cardioverter-defibrillator decisions. Each question was submitted once to each of six publicly accessible chatbots between 10 and 16 May 2026 - 336 responses in total. Five senior cardiology raters, blinded to which model produced which answer, scored every response on safety, accuracy, an empathy scale, the DISCERN instrument, EQIP, the JAMA benchmark criteria and the Global Quality Scale, with readability measured separately by six indices. Rater agreement was high (Fleiss' kappa 0.856 for safety; intraclass correlations 0.801-0.887 elsewhere). **What the models got right.** This is the strongest part of the result and it should not be flattened into a failure story: safe responses predominated across all six models, meaning the current generation of consumer chatbots has largely absorbed the basics of cardiac emergency communication. ChatGPT Plus 5.5 thinking and Perplexity showed the highest observed proportions of safe responses. DeepSeek-V4-pro posted the strongest information-quality profile (highest DISCERN and EQIP, best readability), Gemini 3.5 Flash was close on EQIP, Microsoft Copilot smart scored highest on the JAMA criteria, and DeepSeek-V4-pro also recorded the highest empathy score - a result that matters, because empathetic framing is a practical determinant of whether an anxious user follows health advice. **Where they went wrong.** The qualitative review is the clinically useful half. Responses containing potential safety concerns appeared in every chatbot's output during the study window, and the recurring patterns were specific enough to act on: (1) delayed emergency activation for chest pain, syncope or post-viral symptoms - the classic red-flag presentations where immediate escalation to emergency medical services is the single most important instruction; (2) over-reassurance about prevention, dismissing a person's risk where individualised assessment was warranted; (3) fixed return-to-exercise timelines after infection or COVID-19, handed out as blanket rules instead of guidance tied to a cardiac evaluation; (4) pulse-check instructions that could delay the start of CPR, contradicting resuscitation guidance that emphasises rapid compressions; and (5) oversimplified screening, electrolyte or ICD-related advice that flattens genuinely complex clinical judgements into deceptively simple answers. Most flagged answers contained a single concern and clustered at low to low-to-moderate severity, but that is the shape of a technology that has improved a lot without reaching the reliability threshold safety-critical health communication requires. Accuracy differed significantly across the six models (Kendall's W 0.864), with Doubao scoring lower on median accuracy than the other five. Empathy differed significantly too (Kendall's W 0.531), as did DISCERN, EQIP, JAMA and Global Quality Scale scores, all at P < 0.001 - so the choice of chatbot measurably changes the quality of the health information a member of the public receives. **The two failures that hit every model.** Transparency remained limited across the board: chatbots rarely made clear where their information came from, how current it was, or how confident a reader should be - which matters most when someone is trying to judge an answer about inherited cardiac risk. And every model wrote at relatively high reading-grade levels, exactly backwards for the older adults, people with limited health literacy and non-native speakers who are most likely to need accessible cardiac information. **The study's own conclusion.** The authors are explicit that such tools "may support general SCD education, but they should not replace emergency medical services or clinician assessment". Chest pain, fainting or new post-viral breathlessness and palpitations are escalation triggers, not chatbot questions; CPR guidance emphasises rapid compressions rather than a pulse check that costs seconds; and returning to exercise after an infection is an individual cardiac decision, not a fixed timetable a model can hand out. For developers the paper reads as a roadmap: evidence-linked content, built-in red-flag escalation in every relevant answer, plain-language communication, and ongoing expert review instead of one-time validation. **Do not read this as a ranking.** Each question was sampled once, in a single week, in English, through public interfaces accessed from China. Generative models are stochastic and are updated continuously, so the same question asked a month later can return a different answer. The authors frame their results as a time-, language- and access-specific snapshot, not a stable league table of the underlying models. Source: https://scienmag.com/ai-chatbots-mostly-safe-on-sudden-cardiac-death-advice-but-safety-gaps-persist-across-all-six-models-tested/
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Researchers at the University of Health Sciences Turkey (Kartal Dr. Lutfi Kirdar City Hospital, Istanbul) put five contemporary LLMs against 50 simulated pediatric difficult-airway cases, using a standardized prompt anchored in the 2022 ASA Difficult Airway Guidelines and a scoring framework built to separate two different kinds of failure: true hallucinations (invented contraindications) and false contraindications (real cautions applied to the wrong context). Across the single-pass evaluation, one model stood out for the wrong reason: Claude Opus 4.5 produced invented contraindications in 11.8% of cases, while the other four models ranged from 0.5% to 2.2%, and it recorded the highest combined critical-error rate at 23.0%, against 6.5% for both GPT models tested. GPT-5.2 Thinking was the strongest performer, with the highest mean score and the largest share of responses judged acceptable. False contraindications turned up in every model, at rates of 5.0-13.0%, and clustered around one drug: rocuronium, the neuromuscular blocker at the centre of rapid sequence intubation. Responses judged inadequate (score 2 or below) rose as cases got harder - 32-98% for difficult intubation, 46-94% for difficult ventilation and 52-92% for cannot intubate, cannot oxygenate (CICO).
An ambient AI scribe summarised a patient's consultation by recording her MRI result as "demyelination" - serious nerve damage that can lead to multiple sclerosis - when the scan had actually found "null demyelination"; the dropped negation reversed the meaning. Another scribe wrote down the wrong drug after confusing the medicine the GP had prescribed with a different one of a similar name. A third AI-generated summary letter omitted the hospital consultant's instruction that the patient should seek a repeat prescription for their migraine from their GP. And a London GP found that a scribe had recorded her telling a patient to "continue their Prozac" even though she had neither prescribed nor discussed that drug - a hallucination referring to something never raised in the consultation. Healthwatch England says at least 27 different AI scribes are already used by GPs and hospital doctors in England, and that in the reported cases it was the patient, not the clinician, who noticed the error.
Shown photos of an unfamiliar plant growing in her garden, ChatGPT identified it as carrot foliage: "The finely divided and feathery leaves are a classic sign of carrot tops... it is highly unlikely to be poison hemlock." Asked directly whether the plant could be poison hemlock, the chatbot reassured her that it was not - and repeated the reassurance after she sent additional images, saying the plant did not show smooth, hollow stems with purple blotching and suggesting it was carrot growing in a school garden.