AI explanation errors became harder to judge in the very reasoning format people said they preferred. In a controlled study of 50 participants, researchers compared six ways of showing reasoning from GPT-5. The most popular structured format produced a 19.2% false-alarm rate and 57.1% error-localisation accuracy, while a simpler chain-of-thought format produced a 7.7% false-alarm rate and 95.5% localisation accuracy. The preprint, submitted on 8 September 2026, argues that explanation style can change how well humans supervise model outputs.
The result challenges a comfortable assumption. More structure can feel more transparent. A plan, decomposition or polished sequence of steps looks easier to inspect than a plain answer. Yet the study found a gap between what people liked and what helped them identify mistakes accurately.
Why AI Explanation Errors Can Hide Inside a Better-Looking Format
The researchers generated responses to tasks drawn from GSM8K, HotPotQA and Big-Bench Hard, then deliberately inserted controlled errors so that they could measure whether people detected and located them. They retained 27 problems that GPT-5 originally answered correctly and created errored versions for the human evaluation. This design gave the researchers a known ground truth rather than relying on ambiguous real-world mistakes.
Participants compared six reasoning representations. Some formats exposed a relatively direct reasoning trace, while others organised the answer into a more deliberate plan or structured decomposition. The study found that users tended to prefer the more organised presentation, but preference did not reliably predict oversight quality.
That distinction matters because an explanation has at least two jobs. It can help a reader feel oriented, and it can help a reader test whether the answer is right. Those objectives overlap, but they are not identical. A clean structure may reduce cognitive effort while also making a flawed path feel coherent enough to pass without challenge.
The Most Popular Explanation Was Not the Best Error Detector
The headline result came from comparing the favoured Plan-and-Solve style with a simpler zero-shot chain-of-thought representation. In the study, the structured format triggered more false alarms and made it much harder for users to identify where the actual error occurred. The simpler trace was less popular but much better for localisation.
A false alarm is not harmless. If a reviewer repeatedly flags correct reasoning as wrong, the system becomes harder to use efficiently. Equally, if a reviewer spots that something is wrong but cannot locate the faulty step, correcting the answer becomes slower and less reliable. Human oversight needs both sensitivity and discrimination.
LiveAIWire has seen the same broader problem from another direction. An AI hallucination study found a hidden warning signal inside model behaviour, suggesting that model confidence and presentation can obscure uncertainty rather than reveal it. The new explanation study adds a human-interface layer: even when reasoning is shown, the format can change what people notice.
What This Means for You When Checking an AI Answer
Do not treat a detailed explanation as additional evidence merely because it is detailed. A long, well-labelled chain can contain the same unsupported leap as a short answer. When the decision matters, check the claim that carries the consequence rather than rewarding the response for looking organised.
One useful habit is to separate verification from reading flow. Identify the conclusion, the key intermediate assumption and any external fact the answer depends on. Then test those pieces independently. This reduces the chance that a persuasive narrative structure will carry the reviewer from premise to conclusion without a genuine check.
Another useful habit is to preserve an independent human judgement before seeing the AI’s explanation. LiveAIWire recently covered evidence that human and AI decisions improved when they were formed independently before disagreements were resolved. The principle is relevant here. If the explanation is seen first, it can frame the problem before the reviewer has established a separate view.
Explanation Quality Is Not the Same as Model Quality
The study does not show that GPT-5 became less accurate when it used a structured format. The researchers were testing how humans evaluated pre-generated outputs with deliberately injected mistakes. The dependent variable was oversight performance, not the model’s underlying ability to solve a fresh task.
It also does not prove that simple chain-of-thought should always be shown to users. Reasoning traces can be incomplete, misleading, overly verbose or inappropriate for some product settings. Model providers may also use internal reasoning that is not equivalent to the text shown to a user. The finding is narrower: in this controlled comparison, the presentation people preferred was not the presentation that best supported error checking.
The sample was 50 adults comfortable reading English, recruited through in-class announcements. The tasks came from benchmark-style problems and the outputs were generated in advance. Live workplace decisions, legal analysis, medical information and open-ended research can produce different patterns. The paper is a useful warning, not a universal ranking of explanation interfaces.
AI Products Need to Test Explanations for Oversight, Not Popularity
Product teams often evaluate an explanation by asking whether users understand it or like it. Those measures are reasonable, but this study suggests a third test: does the explanation help people catch the model when it is wrong? A format can score highly on satisfaction while weakening that function.
This is particularly important as AI systems move into tools where a human is formally “in the loop”. A reviewer who sees a polished rationale may technically retain authority while practically becoming easier to persuade. Human control is meaningful only if the interface helps the person recognise when intervention is necessary.
LiveAIWire’s report on an AI paper checker that surfaced hundreds of confirmed mistakes showed the value of tools designed around falsification rather than presentation. The same design philosophy can apply to explanations. Instead of merely telling a smooth story, an interface could highlight assumptions, expose uncertainty, invite counterexamples or make disputed steps easier to test.
Better Human-AI Teamwork May Need Productive Friction
Good interface design usually removes friction. Oversight may be one of the exceptions. If every step is rendered into a highly fluent narrative, the reviewer may move too quickly from reading to acceptance. A small amount of structure that forces comparison, evidence checking or independent judgement can be useful precisely because it interrupts that momentum.
That does not mean making AI deliberately confusing. It means optimising for the real task. If the task is brainstorming, a friendly structured explanation may be ideal. If the task is auditing a calculation, assessing a safety decision or reviewing a consequential recommendation, the interface should be evaluated by how often people catch the right error without inventing new ones.
The broader lesson is that transparency is not measured by word count. An AI can reveal more text and still leave the human less able to judge it. The useful explanation is not necessarily the one that feels most impressive. It is the one that helps the reader know when to trust, when to question and exactly where to look when something goes wrong.
Verification Interfaces Should Show Where a Claim Can Break
The study also points towards a better way to think about explainability. A useful oversight interface should not only organise a model’s reasoning; it should expose the places where a reviewer can test it. That could mean separating factual premises from calculations, marking which statements depend on external evidence and making it easy to compare an intermediate step with the final conclusion.
This matters because explanation length and explanation usefulness can diverge. A reviewer faced with several screens of polished reasoning may spend more effort following the presentation than challenging its assumptions. A shorter trace can sometimes make the decisive transition easier to see, which is consistent with the study’s localisation result even though the experiment does not identify one universal interface design.
The finding also complements LiveAIWire’s coverage of human-AI teamwork. Collaboration performs best when the human contributes something genuinely independent rather than merely endorsing the system. An explanation that makes disagreement easier can therefore be more valuable than one that makes agreement feel comfortable.
For reviewers, that is a useful design test: if an explanation makes the answer feel clearer but does not make the decisive claim easier to challenge, it may be improving presentation more than oversight.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
