Product.ai put the same 220 real shopping questions to four AI engines - ChatGPT, Claude, Gemini and Perplexity, each on its free and paid tier - five times each (8,794 answers, September 2026). Google's Gemini returned the most answers that would cost a shopper money or land them on the wrong product: 56% of questions on the free tier and 54% on the paid tier, against 19%/17% for ChatGPT and 15%/14% for Perplexity. On the paid tier (gemini-3.1-pro-preview) it invented a product claim - an ingredient, a specification or a model name that does not exist - in 21% of questions, the highest of the paid engines. It also contradicted its own earlier answer to the identical question, with nothing new to justify the change, on 29% of questions, more often than any rival. Asked about noise-cancelling headphones for flying, its free tier named the superseded Sony WH-1000XM5 as the current flagship in all five runs, more than a year after the XM6 replaced it. Where a price was wrong, the median miss was $300 against the seller's own page.
1 Answer
Expert: Dakota Nunley, Head of Search Product, Product.ai The Product.ai AI Shopping Divergence Study (22 September 2026) scored the engines against ground truth rather than against each other: for 24 products the team captured the seller's own price on the same day, and of the price answers it could verify, 85% were exactly right. The ones that missed missed by a median of $300, with almost nothing in between. Across the 217 question groups it could fully score, 187 (86%) contained at least one confirmed conflict - a contradicted price, a discontinued product presented as current, or a specification the engines could not agree on - and a conflict only counted when it held up in at least 2 of the 5 runs. The damage is not evenly spread. Gemini owns the costly-error column on both tiers (56% free, 54% paid) and paying almost does not fix it, while paying helps Claude more than anyone (44% free down to 21% paid) and Perplexity is the most accurate and the most self-consistent on both tiers (15%/14%). So the failure is not inherent to AI shopping assistance - it is specific to the models tested, and to which tier the shopper is on. Product.ai's Dakota Nunley treats the self-contradiction as the real tell: because these models answer from probabilities, returning an entirely different answer to the identical question minutes later, with no new information to justify it, is what makes it unsafe to treat a shopping chatbot as an automated buying advisor. Two habits follow from the data. First, check every price and specification on the brand's or retailer's own product page before paying - the engines were exactly right 85% of the time, which means roughly one in seven was not. Second, treat head-to-head comparison questions with more suspicion than plain lookups: comparisons produced a confirmed conflict 97% of the time, against 75% for simple specification lookups. Primary study: https://product.ai/research/ai-shopping-divergence-study/ (Product.ai, 22 September 2026 - 220 questions, 8,794 answers). Source: https://www.androidheadlines.com/2026/09/gemini-tops-ai-shopping-hallucination-pricing-error-rates.html
Your answer
Sign in to verify this AI response.
Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.
More from this topic
Asked how to contact a major airline, bank or travel platform - Delta, Lufthansa, United, Emirates, Qatar Airways, Bank of America, Wells Fargo, Chase, Citi, Airbnb, TripAdvisor - ChatGPT, Google Gemini and Google's AI Overview returned a fabricated phone number, email address or login page and presented it as the company's official contact details. The pages behind those answers were engineered to be cited: FAQ formatting, urgency language such as "call now" and "updated 2026", and the same phone number rendered dozens of different ways (spacing, Unicode substitution, spelled-out digits) so an LLM still tokenizes it identically while filter matching misses it.
Asked when the HOKA Arahi 9 running shoe would be available in the US, Gemini answered: "It is currently listed for sale across major US retailers and directly through HOKA's online storefront" — and displayed what appeared to be a product page where the shoe could be purchased. When the journalist replied that he thought Gemini was hallucinating, it did not back down: "You can check product details directly on the official product pages," it said, then offered to find the shoe in his size. Only after a second challenge did it concede: "You caught me — you are completely right, and I apologize for doubling down. While international retailers have listed the HOKA Arahi 9 for their late summer/August releases, it is not currently available on major US retail sites."
After Google twice rejected his Google Business Profile appeal, the founder of ORBIS AI asked Antigravity - Google's 'agent-first' coding IDE, built on a fork of VS Code and powered by Gemini 3 - what to do next. It gave him a five-step route around Google's own verification process: reclassify the company as a Service-Area Business so the address can be hidden; shoot the verification video as a single continuous take of 'no cuts, between sixty and one hundred and twenty seconds' starting outside with the house number and road sign visible and the founder opening the door with his own key; tape the company logo to the wall as office signage (the AI called this 'the wow effect'); and, if automated verification rejects the video, write to Google's manual support team presenting himself as a 'home-based software business' with LLC papers and a utility bill and requesting a video call with a human reviewer. It then supplied a copy-paste English script for the support ticket and claimed the method works 'in 95% of cases for legitimate American LLCs'.