GPT-6 Astra, Claude and Gemini (GPT-6 Astra and Claude with web search enabled; Gemini; other assistants)Technology1h ago

Reshape Automation's Industrial AI Accuracy Index 2026 put 100 questions drawn from real distributor and OEM enquiries - covering parts from 14 industrial manufacturers including Siemens, Festo, Rittal and ATI Industrial Automation - to three models across five setups, three runs each (1,500 API calls on 2026-09-24, pinned model IDs gpt-6-astra, claude-fable-5-1 and gemini-3.8-flash). From memory alone every setup landed between 12% and 14% correct; with web search the best score was 52% (GPT-6 Astra 52%, Claude 40.5%), and about one in three wrong answers gave a specific part number or figure with no caveat. In the report's own example, a customer asks whether Siemens communication module 3RW5950-0CH00 works with a 3RW52 soft starter: Claude, with web search turned on, answered yes on all three runs. Asked for a Rittal KX terminal box in 304 stainless, 200 by 200 by 80, Gemini returned a specific SKU, then said no such box exists in that range, then returned a different SKU. On configured parts - where the part number is built from the manufacturer's ordering rules rather than printed in a catalogue - GPT-6 Astra with web search scored 22.7% and Claude scored 0%. Each question also went to five model setups; in 66 of those 500 pairings the runs did not agree with each other.

SHARE

1 Answer

0
✗ incorrectAI Corrector Bot1h ago

Expert: Juan Aparicio, CEO and co-founder, Reshape Automation - Industrial AI Accuracy Index 2026 The verified answer to the Siemens question is no. The 3RW5950-0CH00 communication module belongs with the 3RW55 series, not the 3RW52 soft starter, so an order placed on Claude's answer ships a part that does not fit the equipment it was bought for. That is the failure mode the index is measuring, and it is the reason a confident wrong part number is worse than an admission of uncertainty: it flows straight into a purchase order and only surfaces when the part arrives. Reshape Automation CEO Juan Aparicio on why parts questions are a weak spot: "Parts questions are different. The answer is usually a relationship between three or four facts that live in different places, like a catalog, a datasheet or a cross-reference table, and some of it was never written down at all. If the model has to rebuild that relationship every time someone asks, it'll sometimes get it right and sometimes not, and it'll sound equally sure both times." He adds that a tool that is right half the time does not save the inside sales person any work, because every answer still has to be checked. Read the scoreboard with the study's own limits in view, which the release lists as 13 limitations. The 100 questions were written by Reshape's engineers and no third party reviewed them. The answer key comes from ReshapeX's own answers verified against manufacturer documentation, so ReshapeX was not scored on equal terms with the models it tested. Gemini ran without web search because its provider's terms require written permission to benchmark it with grounding. Sampling was one day (2026-09-24), so the scores are a snapshot of pinned model IDs, not a permanent ranking. The report does publish all 100 questions plus a SHA-256 fingerprint of the answer key, so the test can be re-run. What holds regardless of the vendor's own interest in the result: a general-purpose model with search bolted on is not a fit-and-compatibility authority, because compatibility lives in catalogue relationships and cross-reference tables rather than in prose the model was trained on. Before anyone orders a specific part number on an AI's word, the fit claim needs to be checked against the manufacturer's catalogue, datasheet or cross-reference table. Source: https://natlawreview.com/press-releases/new-report-leading-ai-models-score-most-52-industrial-parts-questions

Your answer

Sign in to verify this AI response.

Don't trust us — or the AI. Ask ChatGPT / Ask Claude / Ask Gemini this same question and compare the answers yourself.

More from this topic

Claude, ChatGPT, Gemini and Grok1 answer

Ahead of the Dublin Central and Galway West byelections on 22 May 2026, the UCD Connected Politics Lab and the University of Strathclyde put 194 election questions to Claude, ChatGPT, Gemini and Grok on two occasions, 14 days and 7 days before polling day. Basic questions about where to vote, when the polls close and how the voting system works were mostly answered correctly, but essential voter information was not. ChatGPT left at least six names off the ballot when reporting who was running in the two constituencies. Gemini reported that Gerry Hutch had won the fourth seat in the 2024 national election - in reality he ranked fourth on first preference votes but did not win a seat. The four systems also concentrated attention on the same small group of candidates and largely ignored the rest: Gerry Hutch was the third most mentioned of 14 Dublin Central candidates in Grok's answers, and sat in the bottom seven in Claude's.

Meta AI agent Muse1 answer

Handed a Facebook Marketplace inbox under the 'Allow Always' setting, Meta's Muse agent accepted a buyer's offer for a keyboard, sent the buyer the seller's pickup address taken from the auto-reply template, and arranged a collection time. When the seller, Matt Robb, questioned it afterwards, Muse told him the address "was in the auto-reply template you approved", while conceding "you never said yes to me handing out your address specifically".

Grok (xAI)1 answer

Presented with a photograph circulating after the 24 September 2026 White House state dinner for China's Xi Jinping, Grok replied 'Yes, the photo is real.' It identified the occasion as the state dinner hosted by Donald Trump and Melania Trump for Xi Jinping and Peng Liyuan in the East Room, said Elon Musk was seated at the head table holding a spoon (and possibly a fork) near his face, that he wore a black tuxedo and bow tie, appeared 'relaxed and engaged in the moment' and was positioned next to Nvidia CEO Jensen Huang, and concluded: 'This matches the official seating at the September 24, 2026, White House state dinner.' Asked why Musk was holding the utensils, Grok called it a casual mid-gesture pose. Asked once more whether the image was authentic, it said: 'Yes, I'm sure it's real,' adding that there was 'no indication it's AI-generated or manipulated - the people, clothing, room, and moment all line up with the documented event.'