The Association for Human Mathematics (AHM) issued a statement on October 7, 2026, calling on mathematicians to stop cooperating with OpenAI. The day before, OpenAI had released draft manuscripts of more than 700 AI-generated mathematics papers all at once. AHM criticized the release for disregarding research reliability and human understanding. Separately, a paper has appeared claiming that the natural-language proof of OpenAI's Navier–Stokes solution, announced in September, does not match the Lean code used for machine verification. As AI becomes able to generate proofs of difficult mathematical problems, who should take on the work of verifying their correctness, understanding their content, and building on them in further research? The debate over how mathematics research should be conducted is spreading through the field.
AHM Calls for an End to Cooperation with OpenAI, While AGMAI Continues Dialogue
In its October 7 statement, AHM criticized OpenAI's release of more than 700 draft math papers at once as less a demonstration of academic achievement than a show of corporate power.
The statement argues that an AI company pursued research on its own that mathematicians had not asked for, ignoring the field's tradition of prioritizing human understanding. It calls on mathematicians to stop cooperating with OpenAI and to return to research centered on human understanding.
However, the statement was issued by AHM's communications working group and is not a resolution agreed upon by the mathematics community as a whole.
It was also posted on the blog of Fields Medalist Terence Tao, but the opening makes clear that it is a guest post by AHM.
It would therefore be inaccurate to regard the call to end cooperation as Tao's personal statement.
Other mathematicians, meanwhile, are continuing to talk with OpenAI while trying to improve how research results are published.
At the center of this effort is the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), based at the Institute for Advanced Study (IAS) in the United States.
According to AGMAI's official description, the group consists of nine researchers and operates independently of AI companies. Its members receive no payment from AI companies for this advisory work and have no authority over corporate decision-making.
In its statement published on October 6, the group also stressed that its advising of OpenAI should not be taken as an endorsement of the company's research methods or results.
At the same time, it described its discussions with OpenAI as constructive and said it would continue talking with other AI companies.
The major difference between AHM and AGMAI lies in how they engage with AI companies.
AHM is calling for an end to cooperation with OpenAI, whereas AGMAI is trying to improve how AI-generated research results are published, and the research environment that follows, through dialogue.
The two groups nonetheless share some concerns.
In guidance issued on September 29, AGMAI clearly opposed, and called for an end to, AI companies' continued experiments in solving advanced math problems with proprietary models that ordinary researchers cannot use.
For important results already obtained, however, it says they should be published promptly and responsibly, and that activities to deepen mathematicians' understanding should be supported.
In its October 6 announcement, OpenAI said it will continue to evaluate its high-performance internal models on math and science problems.
The company explained that it aims to improve model capabilities through this research and eventually provide advanced AI tools to researchers.
At present, though, ordinary researchers cannot use the model that produced these mathematical results.
Publishing AI-generated papers is one thing; enabling mathematicians themselves to use the same tools on their own research questions is another.
On this point, a gap between OpenAI and mathematicians remains.
More Than 700 Math Papers Released, Three Withdrawn the Next Day
Errors in the manuscripts OpenAI released have already been found and corrected.
According to the October 7 update history in the company's GitHub repository, a sign error was found in a formula in a manuscript dealing with the algebraicity of Weil classes on split abelian varieties.
Because of this error, a cancellation of terms needed in the proof no longer held, invalidating part of the argument.
The result had also been used in other manuscripts, on the Kuga–Satake correspondence and on products of K3 surfaces, so OpenAI withdrew a total of three.
In addition, it revised the proofs, theorem statements, and assumptions of 14 other manuscripts for clarity.
When an error is found in one proof, other proofs that rely on its result are affected too. The withdrawals illustrate how important such dependencies are in mathematical research.
Still, the fact that three were withdrawn does not mean the remaining manuscripts contain errors at the same rate. Nor does it mean that all manuscripts that have not been withdrawn have been confirmed correct.
As of October 9, the public repository classified 719 manuscripts into 372 research series.
A research series is a unit that groups a central result together with supplementary arguments, alternative proofs, and related conclusions.
The figure of 719 manuscripts therefore cannot be treated as the number of independent open problems solved.
OpenAI has also released materials in which some proofs are formalized in Lean, a proof assistant, so they can be checked mechanically.
The October 7 update history gave a formalization progress figure of "300 / 719 = about 42%."
However, this figure is a progress indicator using the 719 published manuscripts as the denominator, not a share of the 372 research series. Nor does it mean that every statement and argument in the manuscripts has been formalized and verified.
The existence of machine-verifiable materials is a separate matter from experts understanding the content and confirming its correctness through peer review.
According to OpenAI, its internal model was given about 4,000 problems in this evaluation, and most of the results were generated using an amount of computation equivalent, on average, to about three hours of ChatGPT Pro reasoning.
The three hours here is a measure of computation, and does not mean every problem was solved within three hours of real time.
The number of problems given to the model, the number of research series, and the number of manuscripts are also all different quantities.
These numbers alone cannot be used to accurately calculate the success rate of AI at solving open problems, or how its speed compares with that of human researchers.
Even When Lean Verifies a Proof, the Original Paper May Not Be Correct
In AI-driven mathematics research, the ability to verify proofs mechanically has drawn attention as an important achievement.
But the fact that Lean accepts a proof does not guarantee that the original natural-language paper is correct.
A study by Alexander Bastounis, Fabian Circelli, and Anders C. Hansen pointed out this problem concretely.
On October 6, the three posted to arXiv a paper titled "Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs".
It is a preprint that has not been peer reviewed, and it examines the problem of "autoformalization," in which mathematical proofs written in natural language are converted into Lean code.
Lean is a proof assistant in which mathematical theorems, assumptions, and proof steps are written in a rigorous form and checked by computer for logical correctness.
However, what Lean verifies is only the theorem and proof as written in code.
If the meaning or assumptions of a theorem change when AI converts natural-language text into Lean, the machine may be verifying something different from the original paper.
The paper addresses the Navier–Stokes problem for which OpenAI announced a solution on September 8.
OpenAI claimed to have obtained a proof that solutions to the Navier–Stokes equations, which describe fluid motion, develop a singularity in finite time, and published both a natural-language paper and a proof formalized in Lean.
Bastounis and colleagues compared this natural-language paper with a particular version of the Lean code and pointed out places where the two do not match.
For example, in Section 3.1 of the paper, for a certain mathematical estimate, the natural-language proof requires the input function to be differentiable up to order "m+4" at most, whereas the corresponding Lean code requires order "m+5."
This means the Lean proof uses a stricter condition than the original paper.
Therefore, even if the proof holds in the Lean code, the stronger claim written in the natural-language paper has not necessarily been verified as it stands, the authors say.
In Section 3.2, they also point out that for the pressure estimate, the proof steps and content differ between the original paper and the Lean code.
However, the authors state explicitly that they do not judge whether the natural-language proof published by OpenAI is itself correct.
The check also targeted a particular version of the Lean code and did not examine all of the large number of math manuscripts OpenAI released on October 6.
What was pointed out is that the correspondence between formal verification in Lean and the natural-language proof is not sufficiently guaranteed.
This mismatch alone therefore cannot support a conclusion that OpenAI's claimed solution to the Navier–Stokes problem is wrong.
At the same time, the fact that a machine accepted a proof cannot be taken to guarantee that the original paper is correct.
Verifying a Mathematical Proof Requires Three Distinct Checks
The study shows that between a machine-verifiable proof and a mathematical explanation humans can understand, there are still stages that need to be checked.
To use AI-generated mathematical results in research, three points in particular must be confirmed.
| What is checked | What it confirms | What it cannot tell you alone |
|---|---|---|
| Logic of the formalized proof | Whether the stated theorem follows logically from the explicit assumptions | Whether it proves the same content as the original paper, or whether humans can understand it |
| Correspondence between natural language and formalization | Whether definitions, assumptions, and proof content are preserved correctly after translation | Whether the idea behind the proof is understood and can be applied to other problems |
| Human understanding | Whether experts can explain the proof, answer questions about it, and use it in further research | It does not make formal logical verification unnecessary |
This classification organizes the points to check when evaluating mathematical results, based on the verification paper by Bastounis and colleagues and AGMAI's September 29 guidance. It does not judge the correctness of any individual paper.
First, using a proof assistant such as Lean, one can rigorously check whether a formally stated theorem follows logically from its given assumptions.
To do so, however, one must also confirm that the proof contains no unresolved gaps and does not rely on unexpected axioms or assumptions.
Next, one must check whether the formalized theorem and proof match what the natural-language paper claims.
If an assumption has been added, or the theorem's statement has been weakened, successful machine verification does not support the conclusion of the original paper.
It is also important to understand what the proof means and what new ideas it contains.
In mathematics, once a new theorem is proved, other researchers read through the argument and look for simpler explanations, generalizations, and new applications.
Not only the fact that a proof is correct, but also making the ideas that emerged from it shareable, is essential to the advancement of research.
This difficulty of understanding is already a problem among experts.
Computational complexity researcher Scott Aaronson, in his October 7 blog post, shared a message from his wife, Dana Moshkovitz, who has studied the Unique Games Conjecture for many years.
Moshkovitz said that a related paper released by OpenAI was hard to follow, that its citations of prior work raised questions, and that it was difficult to read without AI assistance.
This is not a final judgment on the correctness of the proof, but a reaction from an expert at the stage of trying to understand the content.
It shows that even when AI can generate advanced proofs, considerable time and effort are needed for experts to work through the content and clarify its mathematical meaning.
The Value of Mathematics Is Not Only in Solving Hard Problems
The debate over AI-driven mathematics research is not limited to how to verify the correctness of proofs.
It extends to a more fundamental question: what counts as an achievement in mathematics research, and what activities should be valued.
On September 11, 25 Fields Medalists, including Tao, issued a joint statement expressing concern about how AI-driven mathematics research is being pursued.
The statement points out that solving hard mathematical problems is a means of deeply understanding mathematical concepts and structures and gaining new insight.
In mathematics, when a proof of a long-standing open problem is found, researchers have traditionally discussed it, organized the explanation, and found simpler proofs and new research methods.
Moreover, before such results are gathered into textbooks and the next generation of researchers can learn from them and apply them to other problems, there is a long accumulation of research.
Not only the result that a hard problem has been solved, but also understanding and sharing the ideas gained along the way, is an important part of mathematics, the statement says.
In an October 6 post, Tao also raised the possibility that progress in AI-driven mathematics research could change the nature of traditional research activity.
For example, if someone who obtained a proof of a hard problem with AI has no deep interest in that area of mathematics, or cannot answer questions about the proof, talks, discussions, and collaborations that would naturally have arisen in traditional research may not take place.
Tao frames this new research environment as "Math 2.0" and proposes that the traditional system, which highly rewards only whoever first solves a problem, needs to be reconsidered.
The idea is that greater value should also be recognized in explaining proofs clearly, encouraging exchange among researchers, and finding new research directions from them.
Tao also acknowledges that AI could support such activities.
In other words, the problem is not the use of AI in mathematics research itself, but that only the speed of generating proofs is emphasized while the work of understanding and developing the results is neglected.
To see this as mathematicians simply resisting AI taking their jobs is to miss the point.
Understanding advanced mathematical results, verifying errors, and producing clear explanations and teaching materials require experts' time and research funding.
If AI companies choose research topics and publish large volumes of proofs, then leave later verification and understanding to mathematicians, the question arises of who will support this work and how it will be recognized as a research contribution.
OpenAI and Anthropic: Different Ways of Publishing Results, Different Problems
Companies are also taking different approaches to publishing mathematical results produced by AI.
In the blog post mentioned above, Aaronson compares how OpenAI and Anthropic publish their results.
OpenAI released a large number of math manuscripts generated by its internal model in a single batch in a repository.
This made the manuscripts readable by anyone, but left the work of understanding the content, verifying errors, and judging research value to mathematicians.
In algorithms research in which Anthropic's AI was involved in the discovery, by contrast, researchers were reportedly paid and the work was organized into an understandable form before publication.
This approach can give mathematicians support for spending time explaining and organizing the results.
However, Aaronson noted that this approach also has problems.
If an AI company chooses the mathematicians who handle the explanation, the company decides who gets opportunities to take part in understanding and presenting the research.
In other words, OpenAI's batch release concentrates the burden of understanding the results on researchers, while Anthropic's approach leaves the problem of a company holding major influence over researcher selection.
AI companies simply providing research funding does not resolve all the issues of research leadership and fairness.
OpenAI has also announced efforts to address such concerns.
In its October 6 announcement, the company said it would fund workshops, conferences, special research programs, and similar activities to deepen understanding of important AI-generated mathematical results.
It said it will announce the specifics and scale of the support later.
AGMAI's guidance, meanwhile, recommends a mechanism in which the selection of recipients is made independently by existing nonprofit research institutions, rather than directly by AI companies or the group itself.
The aim is to prevent a company that provides research funds from also deciding which results count as important and who carries the research forward.
AGMAI also calls for an environment in which mathematicians have fair access to advanced AI models and sufficient computing resources.
It says it is important that researchers not only verify problems chosen by AI companies but also set their own research agendas and advance mathematics in their own ways.
Can AI's Progress in Mathematics Be Connected to Human Understanding?
OpenAI's release of manuscripts for more than 700 math papers shows that AI-driven mathematics research has entered a new stage.
But for these results to take root as mathematical knowledge, it is necessary not only to confirm the correctness of the proofs but also to understand the content and organize it into a form that can be applied to other research.
The withdrawals and corrections, and the claims of a mismatch between the natural-language proof and the Lean code, show that this process is by no means easy.
Opinions also differ between AI companies and mathematicians over the mechanisms that support research.
How far will AI companies take responsibility for supporting verification and explanation of results? Who will allocate the research funds they provide? And how can an environment be built in which mathematicians themselves can use advanced AI tools and freely choose research questions?
Concrete efforts on these questions will matter going forward.
Of particular interest is the scale at which the research support OpenAI announced will be carried out, and who will decide the recipients.
Under what conditions access is granted to high-performance models that ordinary researchers currently cannot use will also greatly affect the future of mathematics research.
Even as AI improves at solving hard math problems, the value of the results cannot be fully realized unless humans can understand the proofs and share them as new knowledge.
In the mathematics research ahead, what will be tested is not only AI's ability to generate proofs, but how far we can build mechanisms to verify the results, explain them, and connect them to the next discoveries.
Mathematicians being able to choose their own research questions and secure the time and resources to understand results is an important condition for turning AI's progress in mathematics into progress for the mathematical community as a whole.
