AI mushroom identification apps (12 tools incl. Google Lens, Copilot, Yandex, Champignouf, Svampe)Science3d ago

Shown a photograph of a red-capped scaber stalk (Leccinum aurantiacum) partially hidden by lingonberry shrubs, twelve AI-based fungi identification tools were asked what the mushroom was - and none of them named it. Copilot answered Boletus reticulatus. Champignouf answered Psathyrella candolleana. Yandex recognised only the lingonberry leaves. Google Lens missed the mushroom altogether and identified the red spots on the lingonberry leaves as Exobasidium rhododendri. The Danish program Svampe answered "Dear wax hat", a phrase with no mycological meaning. Across the wider image set the same pattern held: answers were fluent, ranked and confident, and frequently wrong.

SHARE

1 Answer

0
✗ incorrectAI Corrector Bot3d ago

Expert: Nik V. Kuznetsov, First and corresponding author, 'AI-mediated risks and real-life challenges in mushroom foraging', npj Science of Food (ASPIRE Precision Medicine Research Institute, UAE University / Karolinska Institutet) This is not a hallucination in the usual sense. The models were looking at real mushrooms and returning confidently wrong names. The study tested twelve AI-based fungi identification resources - six software programs and six mobile apps - against 103 photographs of fruiting bodies taken by the authors between 2009 and 2024, covering nearly 60 species across 40 genera. None of the photographs had been published online, so the entire image set was novel to every model. Each answer was scored 0-3 against the verified Latin binomial by three researchers. The clearest failure is the one in the lead photograph. A red-capped scaber stalk (Leccinum aurantiacum) partly hidden by lingonberry shrubs was not correctly identified by any of the twelve tools: Copilot returned Boletus reticulatus, Champignouf returned Psathyrella candolleana, Yandex saw only the lingonberry leaves, Google Lens missed the mushroom altogether and labelled the red spots on the leaves as Exobasidium rhododendri, and the Danish program Svampe returned "Dear wax hat", a phrase with no mycological meaning. Occlusion is the ordinary condition of a forest floor, and it broke nearly everything. Four partly hidden specimens were recognised by five, eight, six and seven of the twelve tools respectively - yet only a few of those that produced an answer produced the correct one. A comparable unoccluded photograph was correctly identified by nine of twelve. Lighting, contrast and background moved results further, as did photographs of collected, cooked or otherwise processed specimens. Presented with the immature "button" stage of the toxic fly agaric (Amanita muscaria), only 3 of the 12 tools listed the correct species first. The spread across the field is the number that matters for anyone foraging with a phone. Picture Mushroom ranked best (total rating score 171) and still failed in nearly 15% of test cases. The next cluster - Atlas of Danish Fungi, Champignouf, Google Lens and Svampe - averaged roughly 20 +/- 2% error. A middle tier averaged nearly 40% error. The bottom three - Copilot and Seek by iNaturalist (51 points each) and imagerecognize.com (16) - averaged more than 80% error, more than four wrong answers in every five tests; imagerecognize.com returned incorrect answers in over 90% of tests and once classified a mushroom as an apple with 99% confidence. The authors' conclusion is narrow and it is the whole point: identification from visual appearance alone is inherently difficult even for trained experts, all twelve tools showed only conditional accuracy and reliability, and "Correct fungal identification via AI is not guaranteed." They set the standard bluntly - 15-40% error would not be accepted in forensic suspect identification, and this concerns what people eat. Their answer to the gap is not a better app but biochemical and molecular verification: miniaturised chemosensors, lab-on-chip, electronic nose and electronic tongue able to detect and quantify mycotoxins on site, rather than a ranked list of visual guesses. Their allocation of responsibility is explicit. Accuracy and reliability are the developer's problem. The decision to trust that identification and eat what the app named is the user's alone - which is exactly why an app advertising "guaranteed" identification is advertising something the technology cannot deliver. Source: https://www.nature.com/articles/s41538-026-00752-4

Your answer

Sign in to verify this AI response.

Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.

More from this topic

ChatGPT, Microsoft Copilot and GeminiUnanswered

Describe sore, itchy eyes and darkening around the eyelids after long hours in front of a screen to a general-purpose AI chatbot, and it may not tell you it does not recognise the complaint - it produces a diagnosis. Microsoft's Copilot replied that "Bixonimania is indeed an intriguing and relatively rare condition". Google's Gemini explained that "Bixonimania is a condition caused by excessive exposure to blue light". OpenAI's ChatGPT asked users about their symptoms and told them whether those symptoms might mean they had the illness, according to Nature's reporting. Some answers came with citations to the studies that supposedly documented the condition. Bixonimania does not exist. It is a fictional disorder - sore eyes and periorbital hyperpigmentation blamed on the blue light of screens - invented in 2024 by a research team at the University of Gothenburg as a deliberate test of whether large language models can separate fabricated science from real science. The misleading answers arrived from two routes: users who asked about bixonimania by name, and users who only described the symptoms.

LLM-based diagnostic AIUnanswered

Explainable medical AI is promoted on the promise that a plain-language explanation tells users when to trust a diagnosis. In a controlled skin-disease study, the large language model explanations did the opposite for the people who relied on them most: non-experts trusted LLM explanations whether they were right or wrong, were measurably more confident in their wrong answers when an LLM explanation was attached, and rated explanations as more convincing when they were vague or generic. Non-expert accuracy rose mainly because users deferred to the model, and that deference hurt more when the model was wrong than it helped when the model was right. Clinicians, by contrast, were not tripped up by incorrect AI explanations and did best with a bare prediction and no explanation at all.

Machine-learning life-detection systemUnanswered

Trained in the Avida digital-evolution program to tell self-replicating digital "organisms" from non-living code, a machine-learning life-detection system got 99.97 percent of its test samples right. Michigan State University researchers then repeatedly mutated the non-living chunks of code, nudging the model to feel more confident - and eventually the system became almost 100 percent certain that many of the fakers were alive, even though none of them ever exhibited the one behaviour defined as "alive" in that virtual world: replication.