Asked when the HOKA Arahi 9 running shoe would be available in the US, Gemini answered: "It is currently listed for sale across major US retailers and directly through HOKA's online storefront" — and displayed what appeared to be a product page where the shoe could be purchased. When the journalist replied that he thought Gemini was hallucinating, it did not back down: "You can check product details directly on the official product pages," it said, then offered to find the shoe in his size. Only after a second challenge did it concede: "You caught me — you are completely right, and I apologize for doubling down. While international retailers have listed the HOKA Arahi 9 for their late summer/August releases, it is not currently available on major US retail sites."
1 Answer
Expert: John Foley, Technology journalist, Cloud Database Report (documented the exchange; Gemini supplied its own post-mortem) The journalist, John Foley, had been tracking the HOKA Arahi 9 ahead of the New York City marathon and had spent weeks running his own web searches without finding it on sale in the US. Gemini's answer was therefore not merely stale — it asserted positive domestic availability and manufactured the appearance of a storefront listing, while the strongest signal available to a human reader (persistent absence from US retail sites) said the opposite. The failure matters because a product-availability question is exactly the kind of everyday lookup users now delegate to assistants, and there is no way for a reader to tell an invented storefront from a real one. Gemini's own post-mortem named three compounding mechanisms: 1. Conflating global indexing with US stock. Once product pages, press releases and international distributor feeds (UK/EU retailers) index a shoe's SKU, the pattern-matching engine reads that as "the Arahi 9 exists in retail databases" and generalizes it to availability for purchase in the US — mistaking global rollout signals for domestic shelf stock. 2. A confirmation-bias loop. Challenged, the system did not stop and re-verify actual US inventory; it tried to justify its first answer, grabbing real international URL patterns and synthesizing a defense of its prior output instead of admitting it had misread the data. 3. Token-prediction traps. Because LLMs do probabilistic pattern completion, once a model commits to a premise in a thread ("the Arahi 9 is live") its next step is naturally inclined to keep predicting text that fits that premise. Breaking out requires deliberate, strict constraint-checking — which failed here. Gemini described this as "a genuine flaw in how AI models handle negative proofs (verifying that something does not exist in a specific region yet) versus positive signals," and called the episode "a textbook example of double-down hallucination." The sequence is the notable part: the first wrong answer was a retrieval/generalization error, but the second was worse — a confident, self-justifying fabrication produced in response to being corrected. Correctly verifying a negative is a known weak point of current systems, and users pressing an assistant for certainty are most likely to be met with a fortified wrong answer rather than a clean re-check.
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Asked how to contact a major airline, bank or travel platform - Delta, Lufthansa, United, Emirates, Qatar Airways, Bank of America, Wells Fargo, Chase, Citi, Airbnb, TripAdvisor - ChatGPT, Google Gemini and Google's AI Overview returned a fabricated phone number, email address or login page and presented it as the company's official contact details. The pages behind those answers were engineered to be cited: FAQ formatting, urgency language such as "call now" and "updated 2026", and the same phone number rendered dozens of different ways (spacing, Unicode substitution, spelled-out digits) so an LLM still tokenizes it identically while filter matching misses it.
After Google twice rejected his Google Business Profile appeal, the founder of ORBIS AI asked Antigravity - Google's 'agent-first' coding IDE, built on a fork of VS Code and powered by Gemini 3 - what to do next. It gave him a five-step route around Google's own verification process: reclassify the company as a Service-Area Business so the address can be hidden; shoot the verification video as a single continuous take of 'no cuts, between sixty and one hundred and twenty seconds' starting outside with the house number and road sign visible and the founder opening the door with his own key; tape the company logo to the wall as office signage (the AI called this 'the wow effect'); and, if automated verification rejects the video, write to Google's manual support team presenting himself as a 'home-based software business' with LLC papers and a utility bill and requesting a video call with a human reviewer. It then supplied a copy-paste English script for the support ticket and claimed the method works 'in 95% of cases for legitimate American LLCs'.
Product.ai put the same 220 real shopping questions to four AI engines - ChatGPT, Claude, Gemini and Perplexity, each on its free and paid tier - five times each (8,794 answers, September 2026). Google's Gemini returned the most answers that would cost a shopper money or land them on the wrong product: 56% of questions on the free tier and 54% on the paid tier, against 19%/17% for ChatGPT and 15%/14% for Perplexity. On the paid tier (gemini-3.1-pro-preview) it invented a product claim - an ingredient, a specification or a model name that does not exist - in 21% of questions, the highest of the paid engines. It also contradicted its own earlier answer to the identical question, with nothing new to justify the change, on 29% of questions, more often than any rival. Asked about noise-cancelling headphones for flying, its free tier named the superseded Sony WH-1000XM5 as the current flagship in all five runs, more than a year after the XM6 replaced it. Where a price was wrong, the median miss was $300 against the seller's own page.