In November 1859, a mere nine-page paper appeared in the monthly report of the Berlin Royal Academy of Sciences. It was Bernhard Riemann's "On the Number of Primes Less Than a Given Magnitude" (Ueber die Anzahl der Primzahlen unter einer gegebenen Grösse). In this paper, Riemann wrote that it is "very likely" (sehr wahrscheinlich) that all non-trivial zeros of the zeta function lie on the line where the real part equals in the complex plane—the critical line. He also noted that he had set aside the rigorous proof "after some fleeting vain attempts." This single remark gave birth to the most famous unsolved problem in the history of mathematics.
The Riemann Hypothesis carries such weight because the location of the zeros directly governs the distribution of prime numbers. The zeros of the zeta function function like a "spectrum" that describes how primes are scattered among the natural numbers. If all zeros lie on the critical line, the distribution of primes becomes as "regular" as possible. Conversely, if a zero exists off the critical line, unpredictable fluctuations arise in the distribution of primes.
In 2000, the Clay Mathematics Institute designated this problem as one of its Millennium Prize Problems, offering a $1 million reward for its resolution. 166 years later, the prize remains unclaimed.
Why the "41% Wall" Didn't Move for 50 Years
Even without proving the Riemann Hypothesis itself, there is an approach that establishes a lower bound: showing that at least some percentage of the zeros lie on the critical line. Research gradually raising this lower bound has been a steady endeavor continuing since the mid-20th century.
In 1942, Selberg first proved that a "positive proportion" of zeros lie on the critical line, but did not provide a specific figure. In 1974, Levinson invented a new technique (the mollifier method) and raised the lower bound to (approximately 33.3%). In 1989, Conrey applied the theory of Kloosterman sums to reach (40%). Subsequently, Bui, Conrey, and Young pushed the bound to 41.05% in 2011, and Pratt, Robles, Zaharescu, and Zeindler pushed it to 41.7% in 2020.
| Year | Mathematician(s) | Lower bound | Method |
|---|---|---|---|
| 1942 | Selberg | > 0% (no specific figure) | Introduction of mollifiers |
| 1974 | Levinson | ≥ 33.3% | Mollifier method |
| 1989 | Conrey | ≥ 40% | Mollifier method + Kloosterman sums |
| 2011 | Bui, Conrey, Young | ≥ 41.05% | Two-piece mollifier |
| 2020 | Pratt, Robles, Zaharescu, Zeindler | ≥ 41.7% | Improved mollifier method |
| 2026 | Claude (Anthropic) | ≥ 67.2% | Weil's quadratic form + pair correlation |
As this table shows, over the 46 years from 1974 to 2020, the lower bound moved from 33.3% to 41.7%—an average of only about 0.18 percentage points per year. Within the framework of the mollifier method, the low-41% range was effectively a ceiling.
Meanwhile, a different technique called pair correlation, introduced by Montgomery in 1973, could show—assuming RH—that at least (approximately 66.7%) of zeros are simple zeros. In 2020, Chirre, Gonçalves, and de Laat reached 67.9% using this method. However, all of these results assume RH and cannot be used for an unconditional lower bound.
Herein lay a structural impasse. The pair correlation method, which produces strong results, presupposes RH, while the mollifier method, which can be used unconditionally, plateaus in the low-41% range. A deep gulf lay between the two techniques.
What Bridged the Gap: "Recombining Existing Research"
Between 2023 and 2025, four mathematicians—Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh—published a series of papers removing the RH assumption from Montgomery's pair correlation method. In their 2024 paper (published in Acta Arithmetica), they proved a theorem handling the pair correlation of zeros without assuming RH, and showed that under a condition weaker than RH—namely, that zeros lie within a narrow band near the critical line—more than 61.7% are simple zeros. In a 2025 arXiv preprint, they proved that under a similar condition, the pair correlation method yields more than 67.25% of zeros existing on the critical line.
What Claude discovered was a path connecting these Baluyot et al. results with a paper Bombieri published in 2000 on Weil's quadratic form.
Bombieri's paper studied the quadratic form associated with Weil's explicit formula (an identity linking primes and the zeros of the zeta function). This quadratic form is positive semi-definite if and only if RH holds. Bombieri showed that if RH holds except for finitely many exceptions, the number of negative eigenvalues equals exactly half the number of zeros off the critical line.
According to Anthropic's technical explanation, the core of Claude's approach lies in the following point: constructing a function space with the quadratic form induced by Weil, and simultaneously treating the positive-definite subspace arising from zeros on the critical line and the negative-definite subspace arising from zeros off the critical line. On top of this, an inequality concerning the rank of the quadratic form is derived from information on the first and second moments. Rather than separating the positive-definite and negative-definite parts, treating the quadratic form as a whole in non-diagonal form was the key to drawing a conclusion from the combination of prior results.
As Anthropic itself acknowledges, there is no prospect that the technique used here will lead to a proof of the Riemann Hypothesis. But the fact that the lower bound, stagnant in the low-41% range, was raised to 67.2% shows that discoveries still remain in how existing mathematical tools can be "connected."
650 Failures and 60 Subagents
The process leading to this result also looks quite different from conventional mathematical research.
Jarred Sumner, an Anthropic staff member who is not a mathematician, instructed Claude: "Take a real stab at the Riemann hypothesis." In the first session, Claude generated and tried 650 ideas, all of which failed. When Sumner urged another attempt, Claude spent a day and a half exploring while coordinating roughly 60 subagents. The subagents together executed 2,400 shell commands and wrote hundreds of Python scripts. They performed thousands of numerical verifications against known zeta zeros and peer-reviewed each other's results.
Throughout this process, Claude itself remained skeptical. Anthropic's blog notes that Claude "initially didn't believe meaningful progress was possible." Sumner's input was mostly encouraging messages like "keep going" and "believe in yourself."
After finding the result, Claude also conducted self-verification. It had multiple subagents review the proof, search for counterexamples, download 54 papers from arXiv to check whether the same result had already been published, and independently re-derive the proof from scratch. It then recommended verification by human number theorists.
In terms of computational resources, the two sessions together consumed 31 million output tokens.
How Far Has Verification Progressed?
A mathematical result becomes "knowledge" only once the community accepts the proof as correct. Verification of Claude's result has, at this point, passed through three stages.
First, Anthropic mathematicians Levent Alpöge and Ralph Furman scrutinized the paper and confirmed how the result relates to prior research. Second, experts in the field—Brian Conrey (who himself proved the 40% lower bound using the mollifier method) and Dan Goldston (a co-author of the Baluyot et al. work)—reviewed the paper in a short period. Third, Claude, working together with staff member Eric Easley, completed a formalization of the result in Lean. This formalization passes the standard verification tool comparator.
Lean is a proof assistant that describes each step of a proof in a form that can be mechanically checked by a computer. The success of formalization guarantees that there is no contradiction in the logical structure of the proof. However, what formalization verifies is "the consistency between the stated theorem and the proof"; judgments such as "are the theorem's premises mathematically natural" or "does the result contradict existing literature" are left to human experts.
There are caveats. This result has not been submitted to an academic journal or undergone peer review. The examination by Conrey and Goldston is described as a "favorable examination on short notice," which differs in nature from formal peer review. Anthropic has published the paper, the formalization, and a concise explanatory note on the proof, but verification by the mathematical community as a whole is still a stage yet to come.
What AI's Mathematical Capability Asks Us
How should this result be positioned? Anthropic's blog states that it is an example of Claude being able to "extend the reach and influence of mathematicians' ideas in new ways." This seems an accurate characterization. Claude did not invent a new mathematical technique; rather, it connected existing tools developed from different directions—those of Baluyot et al. and of Bombieri. It appears that AI compressed the time it might have taken a human mathematician to notice that "combining these two might yield something."
At the same time, this result also reflects the limits of AI's mathematical capability. Claude made no headway on the Riemann Hypothesis itself. All 650 attempts failed, and what was ultimately obtained was merely an improved lower bound on a related problem. Anthropic itself does not believe this technique leads to a proof of RH.
There is another point worth noting. Claude itself initially did not believe it could achieve meaningful progress. Anthropic speculates that this is the result of AI having learned, from training data, about both "the difficulty of unsolved problems" and "the limitations of AI models." The phenomenon of a model underestimating its own capability could become a non-negligible issue in the design of AI systems.
Not a few questions remain. Will this result pass peer review? Has the relationship between the conditions of Baluyot et al.'s pair correlation method (the assumption that zeros lie within a narrow band) and the premises of Claude's result been fully sorted out? Can the combination of Weil's quadratic form and the pair correlation method be applied to other L-functions or analogues of the zeta function? And above all, if this kind of "recombination of existing research" is an area where AI excels, where does the role of human mathematicians shift to?
In a field where the 41% wall did not move even over 50 years, AI achieved an update of more than 25 percentage points in a matter of days. Still, the path to 100% remains out of sight.
