The AI model you use through a chatbot can behave differently from what its benchmark score suggests, even when the underlying model family appears to be the same. A Stanford-led audit of ChatGPT, Claude and Gemini found systematic gaps between results obtained through developer APIs and results obtained through consumer chat interfaces. The September 2026 preprint reported that API evaluations averaged 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test-retest agreement than corresponding interface evaluations.
That matters because benchmark scores are often presented as shorthand for how capable a model is. Buyers compare them, journalists report them and developers use them to decide which system to build on. Yet most ordinary users do not interact with a bare model through an API. They use a product with system instructions, safety policies, memory, tool routing, personalisation and other layers around the model.
The researchers tested the same questions through two routes
The study, also listed by the Stanford Regulation, Evaluation, and Governance Lab, audited seven systems across nine benchmarks covering general capability, social bias and sycophancy. The researchers sent matched prompts through APIs and through the chat interfaces used by consumers, then compared both accuracy and consistency across repeated tests.
The average difference was not huge in every task, but it was systematic enough to challenge a common assumption: that an API benchmark can stand in for the real product experience. For ChatGPT, the paper says the gap between API and interface performance exceeded the API-only difference between GPT 5.3 and GPT 5.4 in the tested setup. In practical terms, changing the access route could matter as much as changing a model generation.
A chatbot is more than the model name
Consumer AI products contain layers that are invisible in a benchmark table. A system prompt can tell the assistant how to behave. A router can send different requests to different model variants. Safety filters can block or rewrite outputs. Memory can add user context. Tool access can allow the assistant to search, calculate or retrieve files. Post-processing can alter the final response before it reaches the screen.
An API evaluation often strips many of those layers away so the tester can measure a more controlled model endpoint. That is useful for scientific comparison, but the result may describe the endpoint rather than the product a customer actually uses. The Stanford team calls this a context-validity gap.
Changing API settings did not fully recreate the interface
The researchers tried varying exposed API controls such as system prompts, sampling settings and reasoning configurations to see whether the consumer-interface behaviour could be reproduced. Those changes affected outputs in some cases but did not reliably eliminate the gap. That suggests the difference cannot always be explained by one obvious hidden prompt.
This is important for organisations that conduct their own model evaluations. A bank, school or software company might benchmark several APIs and then assume the ranking will carry over to staff using the vendors’ chat products. The study says that assumption needs to be tested rather than taken for granted.
It complicates public model comparisons
LiveAIWire has covered how small benchmark differences can dominate model launch coverage, including comparisons between frontier models at different price points and speed improvements in Claude without a token-price increase. Those numbers are useful, but they can become misleading when readers interpret them as a direct score for every way the model is deployed.
A benchmark result answers a specific question under a specific setup. If the interface adds different instructions or tools, the task has changed. That does not make the benchmark fraudulent. It means the scope of the claim has to be narrower: this endpoint achieved this score under these conditions.
Consistency can matter as much as average accuracy
The 2.1 percentage point gap in test-retest agreement is easy to overlook. A system that sometimes gives an excellent answer and sometimes gives a poor one can be harder to use in production than a slightly less accurate system that behaves predictably. Repeated evaluations therefore matter when AI is used in workflows where the same type of request arrives many times.
Interface layers can increase or decrease consistency. A safety policy might force similar answers for sensitive prompts. A router might introduce variation by sending similar requests to different back-end configurations. Memory might make answers more consistent for one user and less comparable across users.
The paper is a preprint, not the last word
The research was posted to arXiv in September 2026 and is described by Stanford as forthcoming at EMNLP 2026. At the time of this article, it should be treated as a preprint rather than a completed peer-reviewed journal record. The authors tested a defined set of systems and benchmarks, not every model, interface or user setting.
AI products also change frequently. A vendor can update routing, memory, safety rules or the underlying model without changing the familiar product name. That means an audit of interface behaviour can become stale faster than a conventional software benchmark.
The practical fix is to test the product you will actually use
For a company buying API access, evaluate the API configuration that will go into production. For a team planning to use a consumer chatbot, test the chatbot under the same account settings, tools and workflows staff will use. If both routes matter, test both. The important thing is to stop treating the model name as the entire system.
This also connects with earlier LiveAIWire reporting on AI bias audits that can disagree on model rankings. Evaluation is not only about choosing the right metric. It is about making sure the thing being measured is actually the thing people will use.
For ordinary users, the result explains a familiar puzzle. A model can receive glowing benchmark coverage and still feel less capable in a particular chat session. That experience does not necessarily mean the benchmark was wrong or the user is imagining the difference. The interface may genuinely be changing how the model behaves.
As AI systems become bundles of models, tools, memory and policy layers, the most meaningful unit of evaluation may increasingly be the whole product. Benchmarking the engine is still useful, but it is no longer enough to tell us how the car will drive.
Benchmarks still matter when their scope is clear
The result is not an argument for abandoning benchmarks. Standard tests remain one of the few ways researchers can compare systems repeatedly under controlled conditions. The problem arises when a score produced through one access route is presented as though it automatically describes every product experience carrying the same model name. A benchmark should identify the model version, interface, tools, settings and date closely enough that another evaluator can understand what was actually tested.
Product teams can apply the same discipline internally. If employees will use a chat interface with memory, retrieval, file tools or safety layers, testing only the bare API leaves part of the deployed system unmeasured. Conversely, an interface test can hide whether a weakness comes from the underlying model or from the surrounding product. Running both creates a diagnostic pair: the API gives a cleaner view of model behaviour, while the product test shows what users are likely to encounter.
Versioning will become increasingly important as AI products change quickly. A model label can remain familiar while prompts, routing, tool policies and other components are altered behind it. Evaluation records that capture those details provide an audit trail when scores move unexpectedly. They also make public comparisons more honest because readers can see whether two results describe the same configuration.
For buyers, the implication is practical. Benchmark rankings can help narrow a field, but a final decision should include representative tasks performed through the actual product that will be deployed. The closer an evaluation is to the real workflow, the more useful it becomes. The Stanford-led preprint gives a quantitative reason to treat the route into the model as part of the system rather than a neutral window onto it.
That also argues for caution when benchmark tables are used in procurement. A one-point lead can look decisive while being smaller than the variation created by access route, prompting or repeated runs. Organisations should therefore record uncertainty and repeatability rather than treating a single score as an immutable property of a model. The more consequential the decision, the more important it is to test the exact deployment configuration.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
