AI & Science

Claude Tried to Solve the Riemann Hypothesis. It Failed, Then Found a 67.2% Result

Claude Riemann Hypothesis research illustrated as an AI system explores the zeta function and reaches a verified 67.2 percent lower bound
Claude Riemann Hypothesis research did not solve the famous conjecture, but produced a new lower-bound result on a related zeta-zero question.

Claude Riemann Hypothesis research has produced a result that is both more interesting and less sensational than “AI solved mathematics”: Anthropic says an unreleased research version of Claude failed to solve the Riemann Hypothesis, but during the attempt raised a longstanding lower bound for the proportion of relevant zeta-function zeros on the critical line from 41.6 percent to 67.2 percent. Anthropic published the result on 10 August 2026 and says two of its mathematicians studied and validated the paper, while outside experts Brian Conrey and Dan Goldston examined it on short notice. The details are in Anthropic’s research report.

The number needs immediate clarification. It does not mean Claude solved 67.2 percent of the Riemann Hypothesis. It is not a confidence score and it does not say that 32.8 percent of the zeros violate the hypothesis. The technical paper proves an unconditional lower bound on a related question, while the paper’s own theorem and limitations leave the Riemann Hypothesis itself unresolved.

Claude Riemann Hypothesis Result: What Actually Happened

An Anthropic staff member challenged the unreleased model to take a serious attempt at one of mathematics’ most famous unsolved problems. Anthropic says the first run generated and tried 650 ideas without finding a successful approach. In a second run, the model spent about a day and a half coordinating roughly 60 subagents, which ran 2,400 shell commands, wrote hundreds of Python scripts and performed thousands of numerical checks. Across two Claude Code sessions, the system generated 31 million output tokens. Anthropic published those details in a separate methodology appendix describing how the search was run.

That workload matters because the result did not emerge from one polished chatbot answer. It came from a research process in which agents proposed ideas, rejected failures, ran code, checked numerical examples and critiqued one another’s work. The remarkable part is the workflow that produced a candidate contribution worth serious examination.

What This Means for You

If you are not a number theorist, the practical significance is not that Claude is about to collect a famous mathematics prize. The stronger interpretation is that a frontier AI system appears capable of doing more than retrieving known mathematics: it can search a large space of approaches, combine established results and produce a new argument that can then be subjected to human and formal verification.

That is the same transition LiveAIWire has been tracking in the race to build agentic AI systems. Once models can coordinate subagents, use tools, write code and critique one another, the unit of capability is no longer a single response. It is the entire research loop around the model.

The Riemann Hypothesis Is Still Unsolved

The Clay Mathematics Institute still lists the Riemann Hypothesis as unsolved. In simplified terms, the conjecture says the non-trivial zeros of the Riemann zeta function all lie on a particular vertical line in the complex plane. Clay says the first 10 trillion relevant solutions have been checked, but checking any finite number is not the same as proving the statement for all of them.

The conjecture matters because the zeta function is closely connected to the distribution of prime numbers. Claude did not provide that proof. Anthropic explicitly says it does not expect the techniques used in this work to lead to one. Anthropic states that limitation directly.

So What Does 67.2 Percent Actually Mean?

Mathematicians have long proved that at least some proportion of the zeta zeros lies on the critical line without assuming the Riemann Hypothesis is true. Anthropic says the previous longstanding lower bound was 41.6 percent and that Claude’s argument raises it to 67.2 percent. The result is a statement about a minimum proportion that can be proved to lie on the line, not a measurement of how much of the Riemann Hypothesis has been solved. Anthropic’s concise mathematical note states a 67.25 percent theorem, which is why the public article rounds the result to 67.2 percent.

A lower bound says at least that proportion has been established. It does not classify every remaining zero. That is why “Claude solved two thirds of the Riemann Hypothesis” would be a compelling headline and a mathematically wrong one.

The Result Was Built on Decades of Human Mathematics

Claude’s argument draws heavily on previous work by Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, together with earlier work by Enrico Bombieri. Anthropic says the model found a way to combine those ingredients to improve the lower bound. The company links directly to the prior research and describes Claude’s result as extending the reach of existing human ideas rather than creating the field from nothing.

That matters for how AI discovery should be described. The strongest story is not that Claude replaced mathematicians. It is that a model navigated a large body of human mathematics and assembled a new route through it. That is closer to an unusually fast research collaborator than an autonomous genius appearing without intellectual ancestors.

Anthropic Did More Than Ask Claude to Check Its Own Homework

After the result emerged, Anthropic says two mathematicians on its staff studied and validated the paper and produced a concise informal note for experts. It also says Brian Conrey and Dan Goldston examined the paper on short notice. Those checks are meaningful, but they should not be described as conventional journal peer review. Anthropic is explicit about who examined the work and how.

That distinction is an important part of trustworthy reporting. Company research can be genuinely valuable while still being first-party material. LiveAIWire applied the same caution when examining Anthropic’s reported Claude watermark behaviour: say who made the claim, what was checked and what remains uncertain, rather than quietly upgrading the evidence into something stronger.

The Formal Proof Changes the Verification Equation

Anthropic also says Claude produced a formally verifiable proof of the result and links to a public repository containing the formalisation. That matters because a proof assistant offers a stricter way to check whether specified mathematical statements follow from specified assumptions than simply trusting a language model’s prose. The public Lean repository makes the formal artefact inspectable.

Formalisation does not automatically prove that every informal interpretation surrounding a result is correct. It does, however, create a stronger verification layer. For AI-assisted mathematics, the combination could become increasingly important: models search rapidly and generate candidate arguments, while proof assistants provide a more rigid logical check.

The Failure Is Part of What Makes the Result Interesting

The failed attempt at the Riemann Hypothesis is not embarrassing context to hide. Anthropic says Claude tried hundreds of ideas that went nowhere before finding a narrower contribution on a related problem. Research itself contains failure, and a system that can surface a useful side result while failing at the headline objective may be more scientifically valuable than one optimised to produce confident-looking final answers. The company describes the lower-bound result as an unintended byproduct of the original challenge.

That is a useful corrective to the way AI science stories are often framed. LiveAIWire’s reporting on AI and drug discovery makes the same distinction between accelerating parts of a research pipeline and magically eliminating the hard stages that remain. Scientific value often appears as compression of search and iteration, not instant replacement of the full process.

The Bigger Question Is Whether This Research Loop Repeats

One result, even a formally checkable one, does not tell us how often Claude can make useful original contributions in mathematics. The next evidence to watch is repeatability across problems, fields and independent research groups. Does the same workflow reliably surface ideas experts missed, or is this a rare success selected from a much larger set of unproductive attempts?

LiveAIWire has also examined how Claude differs from other leading assistants, but research mathematics is a different standard from consumer comparison. A model can be excellent at coding or writing and still be unreliable as a mathematical researcher. The benchmark that matters here is independent reproduction of useful discoveries, not brand reputation.

Claude Did Not Solve the Impossible. It Found Something Better Than a Stunt

Claude Riemann Hypothesis research is compelling precisely because the accurate version survives scrutiny. The famous conjecture remains unsolved. The 67.2 percent figure has a precise technical meaning. Human mathematics supplied the foundation, human experts examined the output, and formal tooling provides an additional checking layer. Those limitations and verification steps are all part of Anthropic’s own account.

If this pattern becomes repeatable, the breakthrough may be less about an AI winning a mathematics prize and more about changing the economics of exploration. A research team could test far more conjectures, computational experiments and proof ideas than before, then concentrate human expertise on the few that survive. Claude failed at the Riemann Hypothesis. What it found on the way may be the more useful signal of where AI-assisted science is heading.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.