AI Research

AGI Consistency Problem: Bold 2030 Bet

AGI consistency problem
AGI consistency problem

By Stuart Kerr, Technology Correspondent, LiveAIWire

The AGI consistency problem is the reason Google’s most advanced AI models can win gold medals at the International Mathematical Olympiad yet still trip over basic high school maths, according to Google DeepMind CEO Demis Hassabis. Speaking on the Google for Developers podcast in August 2025, Hassabis argued that this unevenness, not raw capability, is the real barrier standing between today’s large language models and genuine artificial general intelligence. His framing echoes a phrase Alphabet CEO Sundar Pichai has used for the same phenomenon, artificial jagged intelligence, or AJI, describing systems that spike to superhuman performance in one domain and stumble on tasks a child would find trivial.

Hassabis was explicit that scaling data and compute alone will not fix this. He called for tougher, more targeted testing standards capable of identifying precisely where a model’s reasoning breaks down, rather than benchmarks that reward exceptional performance in narrow, well-defined domains while masking failure on everyday tasks. The distinction matters because it reframes the AGI debate away from a single headline capability score and toward a harder, less flattering question: does the system perform reliably across the full range of situations a genuinely general intelligence would need to handle.

Why Olympiad Gold Doesn’t Mean What It Looks Like

The Olympiad comparison is not incidental. Google DeepMind’s own systems achieved gold-medal-level performance at the International Mathematical Olympiad in 2025, a genuine technical milestone in formal mathematical reasoning. But Olympiad problems are narrow by design, testing deep expertise within a fixed, well-understood problem space. Basic high school maths, by contrast, requires exactly the kind of flexible, context-sensitive reasoning that current models still handle unevenly, sometimes producing an error a distracted ten-year-old would catch. The gap between the two is precisely what Hassabis means by inconsistency, and it is a harder problem to solve than simply building a bigger model.

The Academic Case Against Rushing the Declaration

This is not a new concern, and Hassabis is far from the only person raising it. Cognitive scientist Melanie Mitchell has argued for several years that intelligence is inherently context-dependent, and that AI systems trained on narrow benchmarks remain brittle when faced with genuinely novel situations, a fallacy she attributes to researchers mistaking performance on a specific task for the presence of general, transferable understanding. A separate strand of academic argument goes further, contending that treating a single, monolithic AGI as the field’s north-star goal is itself a mistake, and that AI policy and funding would be better served by prioritising diverse, specialised systems built for specific human needs over one system designed to be generally human-equivalent.

What the Consistency Problem Means for How You Use AI Today

For anyone relying on AI tools in daily work, the practical lesson is straightforward. A model’s performance on a hard, well-known benchmark tells you almost nothing about how reliably it will handle an unfamiliar, everyday version of a similar task. The safest working assumption is that any AI system might fail unpredictably on tasks that look simple, particularly ones involving multi-step logic, careful counting, or context outside its training distribution, precisely the pattern Hassabis describes. Treating AI output as a draft that needs checking, rather than a finished answer from a system assumed to be uniformly capable, remains the most reliable safeguard while this gap persists.

Building Toward Consistency, Not Just Scale

Some architectural approaches are aimed directly at this problem. Mixture-of-reasoners style architectures, which combine multiple specialised reasoning modules under a unified control system rather than relying on one monolithic network, are one of several approaches researchers hope will reduce erratic, task-dependent failures. Whether any single architecture solves the consistency problem outright remains genuinely unresolved, and Hassabis himself has been careful not to claim scale alone will get there. The more likely path, based on his own comments and the wider research literature, is a combination of better testing regimes, architectural experimentation, and a willingness to treat inconsistency as a first-class problem rather than a footnote to otherwise impressive benchmark results.

None of this settles the broader argument over when, or whether, AGI arrives on anything like the timelines industry leaders have floated. What Hassabis’s comments do settle is a narrower but more immediately useful point: judging how close any system is to general intelligence by its best result on its hardest benchmark is a category error. The more honest measure, and the one worth watching as competing labs make their own AGI claims over the next several years, is how consistently a system performs on the ordinary, unglamorous tasks nobody bothers to put in a press release.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, emerging technology, and their impact on business, society, and everyday life. LiveAIWire publishes original AI journalism every weekday at liveaiwire.com.