On October 6, 2026, OpenAI published 722 manuscripts and supporting proof materials on GitHub, all of them mathematical research produced by an unreleased internal AI model. The related papers are grouped into 372 sets of results, and the collection includes research on the zeros connected to the Riemann Hypothesis and a manuscript on the properties of pi.

In September, the company claimed to have solved more than 100 open problems. By releasing the individual conclusions and verification materials in a form outsiders can check, it has given mathematicians more to work with when evaluating the content. Still, publishing manuscripts is a long way from making their proofs understandable to humans and usable in later research.

AD

What do 722 manuscripts and 372 result groups count?

The published openai/math repository contains the paper PDFs along with summaries of the results and materials for mechanically verifying some of the proofs.

The figure 722 is the number of individual manuscripts released. The 372 result groups, by contrast, are categories that bundle a central conclusion with its supporting arguments, the results that follow from it, alternative proofs, and so on. It does not mean that 722 independent open problems were solved.

OpenAI explains that existing math benchmarks had become poor at distinguishing model performance, so it expanded its evaluation to open research problems. The model was given about 4,000 problems, its outputs were organized into related results, and those of some significance were assembled into this collection of manuscripts.

The roughly 4,000 problems, 372 result groups, and 722 manuscripts are therefore different units. Dividing 722 or 372 by 4,000 does not yield an "accuracy rate" or "problem-solving rate" for the model.

Caution is also needed on compute. According to the company, most results were obtained through a common procedure, and the average compute per result is equivalent to having that internal model think for about three hours in ChatGPT Pro.

This is an expression of converted compute. It is not a measurement showing that every problem was solved within three hours of real time, nor a guarantee that entering the same problem into the consumer version of ChatGPT would reproduce the same result.

Two items are exceptions to this common procedure: research on regions where zeta function zeros do not exist, and the Hodge conjecture for abelian varieties with complex multiplication. Some manuscripts dealing with the region where the real part exceeds 11/12 were also edited by humans to make them easier to read.

Average compute alone therefore does not allow a uniform inference about how much work each result required or how much humans were involved.

What has been released is a set of preprints listing OpenAI as the author. They were not published as research that has passed peer review as a whole, and the company itself acknowledges that some results that have not been formalized may contain problems.

What do the Riemann Hypothesis, pi, and Hodge conjecture manuscripts claim?

The September 30 manuscript, titled "The Quasi-Riemann Hypothesis," claims that for all Dirichlet L-functions, including the Riemann zeta function, there are no zeros in the half-plane where the real part of the complex variable exceeds 7/8.

A zero is a point where a function's value is zero. Where the zeros of the Riemann zeta function lie is deeply connected to the regularity of how prime numbers are distributed.

The key point here is the 7/8 boundary. The Riemann Hypothesis itself asserts that all non-trivial zeros lie on the line where the real part is 1/2.

Even if it were proven that there are no zeros to the right of real part 7/8, that would not rule out zeros between 1/2 and 7/8. If the manuscript is correct, it would be an important advance, but it would not solve the Riemann Hypothesis itself.

A more familiar example involving a well-known number is the September 24 manuscript titled "The irrationality exponent of pi is 2."

The irrationality exponent measures how closely an irrational number can be approximated by fractions. It concerns how extremely precise an approximation can be maintained as the denominator grows when pi is approximated by fractions such as 22/7 or 355/113.

Stated concretely, the manuscript's claim is that for any positive ε, if the denominator q is sufficiently large, then whichever integer p is chosen, the approximation error is at least q to the power of −2−ε. In other words, it limits the possibility that extremely precise approximations, relative to the size of the denominator, continue indefinitely.

However, it does not claim that the boundary for what counts as "sufficiently large" can be computed explicitly, nor does it give a lower bound on the error using a single constant common to all fractions.

Another September 30 manuscript claims to prove the rational Hodge conjecture for complex abelian varieties with complex multiplication (CM).

Roughly speaking, the Hodge conjecture asks whether certain structures appearing in geometric spaces can be expressed from parts described by algebraic equations. The manuscript says it holds in any dimension for abelian varieties with CM, but the scope is strictly limited to those. It cannot be taken as a resolution of the Hodge conjecture as a whole, which covers general objects.

The more famous the open problems named, the more necessary it is to check not just the problem names but what is actually claimed to have been proven, and under what objects and conditions.

These manuscripts are subjects for experts to verify going forward. We are not at a stage where one can simply pull out the problem names and treat them as a list of resolved open problems.

AD

What Lean can confirm, and the figure of 162

The formalization catalog, checked on October 7, 2026, lists 162 manuscripts whose main results have been formalized.

That is 22.4% of the 722 published manuscripts. The percentage was calculated by dividing the 162 entries in the formalization catalog by the 722 manuscripts stated in the README, multiplying by 100, and rounding to one decimal place.

However, the 22.4% figure only indicates the share of manuscripts whose main-result formalization materials are registered in the catalog. It is neither the share of manuscripts whose entire text has been formalized nor the share whose verification by independent third parties is complete.

The same catalog lists 185 main theorems, but since one manuscript can contain several theorems, the number of manuscripts and theorems need not match. The catalog's review status is also listed as "unconfirmed," so the existence of formalization materials cannot be regarded as completed review.

Lean is a language and proof assistant that lets a computer check whether each step of a proof for a theorem, with its definitions and assumptions made explicit, follows the rules of logic. It can convert logical leaps that are easy to overlook in prose into a form a machine can verify.

The materials also include settings for Comparator, which checks a submitted formal proof against a predefined challenge.

Under the specified conditions, Comparator confirms that the submitted solution proves the same theorem as the challenge, that it uses only permitted axioms, and that the Lean kernel, the core of its verification, accepts the proof.

This makes it possible to verify not merely that "some Lean code ran," but also exactly what was proven.

Still, the extent of formalization has to be checked theorem by theorem.

For example, the setup for the Quasi-Riemann Hypothesis specifies a theorem stating that the Riemann zeta function does not vanish to the right of real part 7/8.

The existence of this setup does not allow one to conclude that all the L-functions the manuscript addresses, or every explanation in the paper, have been independently formally verified.

Moreover, even when a formalized challenge and its solution correspond correctly, whether the challenge itself accurately expresses the original mathematical problem has to be checked separately.

What a proof adds to existing research, and whether the ideas used in it can be applied to other problems, cannot be known merely because a formal proof passed the kernel. Machine verification of logic and expert mathematical understanding play different roles.

The mathematics community's publication guidelines and the GitHub release

AGMAI, an independent advisory group that discusses mathematics and AI, collected more than 600 responses from the mathematics community and on September 29 released guidelines for publishing AI-generated mathematics. OpenAI says it incorporated AGMAI's advice into how it made this release.

Comparing the guidelines with OpenAI's October 6 announcement lets us separate what has already been done from what remains to be done.

Item compared AGMAI's September 29 guidelines OpenAI's October 6 release
Storage and revision Provide an academic repository not controlled by the AI lab, persistent citation identifiers, and a change history Published on OpenAI's GitHub. Provides citation information for each manuscript and a revision policy that preserves earlier versions; a community-run alternative repository is under consideration
Disclosure of the generation process For each result, state the model name, prompts, reasoning summary, time, and estimated compute cost States that an unreleased internal model was used. Publishes 10 reasoning summaries, average compute, and the number of attempts (about 4,000 problems)
Formalization Formalize proofs wherever possible and attach a catalog and challenges for comparative verification Publishes Lean materials, a catalog, and Comparator settings. Formalization is limited to some manuscripts, with more to be added
Human understanding Fund activities that advance understanding of the content, leaving decisions on who is supported to existing nonprofit institutions Announces funding for workshops, conferences, and special programs to understand the results. Details to be announced

This table matches the four items in AGMAI's guidelines against what OpenAI's announcement and README state explicitly. It does not judge whether all 722 manuscripts comply with the guidelines.

OpenAI has published revision history and formal verification materials, but access to the model that produced the results, and a storage location independent of OpenAI, remain open issues.

Being able to trace results is not the same as mathematicians being able to use the same research tools themselves.

AGMAI's guidelines in fact do not support efforts to have proprietary models, unavailable to the broader research community, solve large numbers of advanced open mathematical problems, and call for such efforts to stop.

At the same time, the guidelines call for prompt publication of whatever results are produced. They also ask AI research organizations not to push the burden of understanding the content onto mathematicians alone after publication.

In its October 6 statement, issued on the day of the release, the group describes this development as significant for mathematics. It also states clearly that the fact it advised OpenAI does not mean it assessed the impact of the published results or endorsed the method by which OpenAI produced them.

AGMAI operates independently of companies and is not an organization representing the mathematics community as a whole. It would not be appropriate to take OpenAI's consultation with AGMAI as mathematicians' endorsement of the results.

AD

Can the published proofs lead to the next research?

OpenAI has pledged funding for conferences and other efforts to support understanding of the results, and says it is working to release the model that produced them in a responsible way.

However, this announcement gives no timing for the model's availability or terms of use, and the specific mechanism for the funding support has yet to be revealed.

One of the concerns AGMAI raised is that future mathematical research must not become skewed toward the work of "humans understanding results produced by AI companies."

Mathematicians themselves need to be able to pose questions, develop their own methods, and freely explore research themes that companies did not choose to showcase the AI's capabilities. As a condition for this, AGMAI points to an environment with fair access to high-performance research tools and sufficient computing resources.

This is also a question of whether the researchers receiving the proofs have the initiative to choose their next research topics.

Even if completed manuscripts can be read, if new questions arising from them cannot be tested with a model of comparable capability, a bias may remain in which directions can be researched. Publishing papers and building an environment in which research can proceed from those results are closely related but not the same.

The focus going forward is whether experts can verify the individual proofs, explain their ideas clearly, and show how they can be applied to other problems.

Whether access to the model and research support will extend to people studying themes other than those OpenAI itself chose will also matter.

If it gets that far, this mass release will be more than a list of AI achievements: it could become knowledge and tools that let mathematicians pursue new research starting from their own questions.