AI developers claim their chatbots are 'safe' for mental-health conversations, based on safety evaluations where hired psychiatrists grade chatbot responses as safe or unsafe.
1 Answer
Expert: Kiana Jafari, Nina Vasan, Stanford Center for AI Safety / Stanford Psychiatry The claim that AI safety testing reliably separates safe from unsafe chatbot responses is not supported. Stanford researchers found that board-certified psychiatrists structurally disagree when grading AI responses to mental-health prompts: three psychiatrists evaluating 360 AI responses produced ratings that, when averaged, matched no single expert's judgment. A poll of 100+ psychiatrists at the American Psychiatric Association annual meeting produced the same near-even split — more experts did not help, because clinicians apply incompatible frameworks (safety-first, engagement-centered, culturally informed). Kiana Jafari (director, Stanford Center for AI Safety): 'It doesn't matter how many experts you have — 3, 10, or 1,000 — when they do not agree, you are not actually getting to the ground truth by averaging their scores.' In high-risk areas like suicidal thoughts, psychosis and eating disorders, 'AI safety is not yet there' (co-author Nina Vasan, Stanford psychiatry). The paper was accepted to ACM FAccT 2026 (arXiv 2601.18061). Source: https://news.stanford.edu/stories/2026/07/study-exposes-major-flaw-in-ai-mental-health-safety-testing
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Generative AI citation tools produce references that look real but do not exist. A 2026 analysis documented 946 hallucinated-citation cases as of February 16, 2026: 647 fabricated citations, 192 misrepresenting facts or precedent, and 107 containing false quotes. These fabricated references appear in legal filings and scholarly papers alike.
Nature analysis suggests tens of thousands of publications from 2025 might include invalid references generated by AI. Computer scientist Guillaume Cabanac (University of Toulouse) received a Google Scholar alert that his paper had been cited in the International Dental Journal — 'I was very surprised to see that I couldn't recognize my own reference.'
AI writing tools are injecting fabricated references into scientific literature. A Lancet audit of nearly 2.5 million PubMed papers found 1 in 277 papers published in the first seven weeks of 2026 cited a paper that did not exist — up from 1 in 458 (2025) and 1 in 2,828 (2023), a 12-fold rise in two years. The sharpest increase began mid-2024, coinciding with the rise of AI writing tools.