Researchers at the University of Health Sciences Turkey (Kartal Dr. Lutfi Kirdar City Hospital, Istanbul) put five contemporary LLMs against 50 simulated pediatric difficult-airway cases, using a standardized prompt anchored in the 2022 ASA Difficult Airway Guidelines and a scoring framework built to separate two different kinds of failure: true hallucinations (invented contraindications) and false contraindications (real cautions applied to the wrong context). Across the single-pass evaluation, one model stood out for the wrong reason: Claude Opus 4.5 produced invented contraindications in 11.8% of cases, while the other four models ranged from 0.5% to 2.2%, and it recorded the highest combined critical-error rate at 23.0%, against 6.5% for both GPT models tested. GPT-5.2 Thinking was the strongest performer, with the highest mean score and the largest share of responses judged acceptable. False contraindications turned up in every model, at rates of 5.0-13.0%, and clustered around one drug: rocuronium, the neuromuscular blocker at the centre of rapid sequence intubation. Responses judged inadequate (score 2 or below) rose as cases got harder - 32-98% for difficult intubation, 46-94% for difficult ventilation and 52-92% for cannot intubate, cannot oxygenate (CICO).
1 Answer
Expert: Merve Bulun Yediyildiz & Irem Durmus, Anesthesiologists, University of Health Sciences Turkey (Kartal Dr. Lutfi Kirdar City Hospital, Istanbul) The point of the study is not that the models scored badly. It is that they fail in characteristically different ways, and the difference decides patient risk. A true hallucination is an invented contraindication - the model says a drug cannot be used in a situation where the guidelines make it the recommended choice. That behaviour was almost entirely confined to Claude Opus 4.5: invented contraindications in 11.8% of the 50 cases, against 0.5-2.2% for the other four models. Claude Opus 4.5 also carried the highest combined critical-error rate at 23.0%, versus 6.5% for both GPT models tested. A false contraindication is subtler: a genuine clinical caution applied to the wrong patient or the wrong moment, so the model withholds an intervention that was appropriate. Every model did this, at 5.0-13.0%, and the errors bunched around rocuronium - the neuromuscular blocker that is central to rapid sequence intubation. A model that wrongly flags rocuronium does not make a harmless stylistic mistake; it steers a trainee away from a potentially lifesaving drug in the exact scenario where seconds decide the outcome. Two other findings matter for anyone tempted to use these tools for exam preparation. First, the failure rate tracks clinical complexity in the wrong direction: inadequacy reached 52-92% on cannot intubate, cannot oxygenate (CICO), the most dire scenario, exactly where accurate guideline-concordant guidance is worth most. Second, paediatric airway work punishes approximation - anatomy, drug dosing and equipment sizing all scale with age and weight, and the therapeutic window for neuromuscular blockers and induction agents is narrow, so a drift at an early node of the ASA decision tree cascades into a wrong endpoint. Human scoring here was not the weak link: two blinded anaesthesiologists rated every response on a modified 0-5 scale, with a quadratic weighted kappa of 0.982, and the model differences were significant (Friedman chi-squared 100.12, p < 0.001). The authors' conclusion is narrow and clear. Current LLMs may hold value as supervised supplementary learning tools, but they should not function as autonomous sources of clinical or educational guidance. A chatbot whose answers must always be checked is a study partner, not a reference. Read as a benchmark lesson rather than a ranking: these questions probe what kind of wrong an answer is, not just whether it is wrong. Study: Yediyildiz, M. B., & Durmus, I. (2026). Mechanism-aware benchmarking of large language models as learning aids in pediatric difficult airway training: true hallucinations vs. false contraindications. BMC Medical Education. https://doi.org/10.1186/s12909-026-10509-y Source: https://scienmag.com/ai-chatbots-flunk-pediatric-airway-emergencies-hallucinations-and-false-warnings-exposed/
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
An ambient AI scribe summarised a patient's consultation by recording her MRI result as "demyelination" - serious nerve damage that can lead to multiple sclerosis - when the scan had actually found "null demyelination"; the dropped negation reversed the meaning. Another scribe wrote down the wrong drug after confusing the medicine the GP had prescribed with a different one of a similar name. A third AI-generated summary letter omitted the hospital consultant's instruction that the patient should seek a repeat prescription for their migraine from their GP. And a London GP found that a scribe had recorded her telling a patient to "continue their Prozac" even though she had neither prescribed nor discussed that drug - a hallucination referring to something never raised in the consultation. Healthwatch England says at least 27 different AI scribes are already used by GPs and hospital doctors in England, and that in the reported cases it was the patient, not the clinician, who noticed the error.
Shown photos of an unfamiliar plant growing in her garden, ChatGPT identified it as carrot foliage: "The finely divided and feathery leaves are a classic sign of carrot tops... it is highly unlikely to be poison hemlock." Asked directly whether the plant could be poison hemlock, the chatbot reassured her that it was not - and repeated the reassurance after she sent additional images, saying the plant did not show smooth, hollow stems with purple blotching and suggesting it was carrot growing in a school garden.
Automated AI claims review is presented as a straightforward upgrade for health, home and auto insurers: it speeds claims processing and reduces errors, so the patient's bill is handled faster and more consistently than a human adjuster could. Adoption is near-universal -- a National Association of Insurance Commissioners 16-state survey found 84 percent of US health insurers already use AI for tasks such as prior authorization, and by 2023 nearly 88 percent of auto insurers were using or planning to use AI for claims -- so an automated decision on a routine claim can be treated as equivalent to a human review of the same file.