OpenAI Releases 722 Mathematical Manuscripts at Once

OpenAI recently released a collection of mathematical research materials generated by unreleased frontier models, comprising 722 manuscripts and covering 372 families of results. It also disclosed selected reasoning summaries, compute estimates, and problem statistics. The true impact remains to be determined once the mathematics community completes its review.
OpenAI Releases 722 Mathematical Manuscripts at Once: AI Problem-Solving Enters the “Pile of Papers” Phase
On October 7, OpenAI announced a rare collection of mathematical research materials: 722 manuscripts covering 372 result families. The materials came from an advanced model that has not yet been officially released. OpenAI said the model had solved or advanced hundreds of long-standing open problems.
The focus of this release was not yet another competition problem answered by a model, but the beginning of AI producing, in volume, mathematical results across fields that require peer review. The materials cover areas including high-dimensional geometry, coding theory, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics. In other words, OpenAI is trying to demonstrate not that “the model can solve math problems,” but that it can continuously generate candidate results worth mathematicians’ attention across multiple research branches.
However, these materials cannot yet be directly equated with “AI has solved 372 difficult mathematical problems.” In mathematics, “solved” has a strict threshold: the proof must be complete, the definitions must be precise, the chain of reasoning cannot rely on model hallucinations, and the result must withstand independent verification by researchers. This release is more like making public a vast “mine of candidate proofs.” The truly time-consuming work that follows will still be for human mathematicians to extract, verify, and organize them one by one.

From Ten Representative Results to 722 Manuscripts
This release did not appear out of nowhere. In August this year, OpenAI had already announced a set of results in mathematics and theoretical computer science, saying that an internal version of its next-generation flagship model, Astra, had achieved ten advances. These results covered problems involving high-dimensional sphere packing, binary and spherical codes, non-sofic groups, the Connes rigidity conjecture, arithmetic circuit complexity, quantum parallel repetition, the closest vector problem, the Ehrhart volume conjecture, and multicolor Ramsey numbers.
Some of these problems are directly related to theoretical computer science and cryptography. For example, the closest vector problem is one of the foundational problems in lattice theory, and its approximate hardness is connected to post-quantum cryptography. Quantum parallel repetition, meanwhile, seeks to extend an important result from classical complexity theory to the setting of quantum games. For developers, this research will not immediately become a new SDK interface, but it may influence the long-term technical boundaries of complexity theory, cryptanalysis, and quantum computing.
The ten results announced in August were more like a capability demonstration: OpenAI selected a small number of representative problems to show that the model could propose proofs, generate paper drafts, and attempt to formalize the arguments as proofs that Lean could verify. This release greatly expands the scope to 722 manuscripts and 372 result families, showing a much stronger tendency toward bulk production.
That is the most significant difference between the two releases. A single impressive result may be an accidental peak of model capability, or it may be a “survivor” left behind after extensive manual screening by a research team. Hundreds of result families, by contrast, suggest that the model has been operating continuously across multiple directions and organizing its outputs into a relatively systematic body of research materials. It is still far from a mature automated research pipeline, but it is no longer merely a laboratory demonstration.
What Process Information Did OpenAI Disclose?
OpenAI placed the materials in a GitHub repository and simultaneously established guidelines for manuscript revisions and citation handling. The company said it would continue looking for publication arrangements hosted by the academic community, while improving the quality of the papers, the completeness of citations, the way mathematical content is explained, and the presentation of results.
The public materials also include summaries of some model reasoning processes, estimates of computing costs, and statistics on the number of problems the model attempted. OpenAI previously disclosed that, for the ten representative results announced in August, the total Token cost of having the model search for solutions was approximately $2,000 based on Sol API rates. The same model was then used to organize the papers and generate formalized proofs in Lean.
In its description of this large-scale release, OpenAI also provided a more intuitive reference point: the computing cost of an ordinary result was roughly equivalent to three hours of continuous thinking by ChatGPT Pro. This statement helps convey the scale of the cost, but it cannot simply be converted into “each theorem costs only a few dollars.”
The reason is simple: the cost of a research task includes more than the manuscript that ultimately remains. The model may need to try numerous incorrect paths, repeatedly revise prompts and context, call formal verification tools, filter out duplicate results, and then have researchers organize and review the material. The publicly disclosed “cost per result” is generally closer to the cost of effective output than to a bill for the complete experimental process.
In addition, summaries of model-generated reasoning processes should not be treated as the “verbatim original thought” in any strict sense. For researchers, what truly matters is a verifiable proof, clear definitions, executable formal checks, and sufficiently complete experimental records—not a seemingly coherent block of explanatory text.
372 Result Families Do Not Mean 372 Grand Prize Problems
“Result families” is a figure that is easy to misinterpret in this coverage.
A mathematical research family may include a central conjecture, multiple parameterized versions, improvements to upper or lower bounds, proofs of special cases, counterexample constructions, or several conclusions derived from the same technical approach. It is not the same concept as “372 completely independent, major problems that had never previously been solved.”
Likewise, “hundreds of unsolved problems” needs to be evaluated in the context of the specific papers. Some results may completely resolve a public problem; some may improve a known upper bound by one step; some may provide a new construction or counterexample; and others may hold only for particular dimensions or parameter ranges. In mathematical research, the value of these results can differ substantially and cannot be judged by quantity alone.
In September, OpenAI said that its model had solved more than one hundred long-standing public problems spanning most branches of mathematics. The public release of 722 manuscripts is clearly an attempt to turn that promotional claim into research materials that can be examined externally. However, increasing the number of materials does not automatically increase the credibility of the conclusions. Instead, it magnifies the issues of review, deduplication, error detection, and attribution.
For the mathematical community, the most important question is not how thick this collection of manuscripts is, but how many can be independently understood and verified without “verbal supplementation” from OpenAI researchers; how many results can be reproduced by other teams; and how many new ideas will genuinely be incorporated into subsequent papers and research programs.
What AGMAI Is Concerned About Goes Beyond Whether the Proofs Are Correct
The Mathematics and AI Advisory Group (AGMAI) issued its first recommendations in late September, calling on AI laboratories to publish mathematical research results promptly through established academic channels and to disclose details such as the model name, prompts, and computing costs.
AGMAI also specifically warned that companies should not package mathematical achievements as model marketing campaigns. The warning is not surprising. AI laboratories are competing for control over the narrative surrounding frontier-model capabilities, and mathematics is one of the most suitable fields for demonstrating “deep reasoning”: its problem boundaries are clear, its answers can be formalized, and its proofs can in principle be checked step by step by peers. An apparently elegant theorem can serve both as technical communications material and as a capability advertisement ahead of a model release.
The problem is that the standards for scientific communication and product marketing are different.
A product launch can emphasize peak performance, best-case examples, and demonstrations. Academic research, by contrast, must explain failed attempts, limitations, related work, author contributions, and paths to reproducibility. If a company publishes only its best results without explaining how many attempts the model actually made, how the prompts were designed, or what key decisions were made by human researchers, it is difficult for outsiders to determine whether the result came from the model itself or from an undisclosed process of manual screening.
That is why AGMAI recommends disclosing the model name and prompts. For developers, prompts are not irrelevant peripheral information. In long-horizon mathematical tasks, how the problem is decomposed, whether the model is allowed to call tools, how context is maintained, and how failed results are fed back can all significantly affect the final output. Without this information, the claim that “the model solved the problem” is difficult for other teams to reproduce.
Lean Formalization: A More Rigorous Quality Threshold
In its earlier representative results, OpenAI said it would have the model formalize mathematical arguments as proofs verifiable by Lean. This is more meaningful than simply publishing a paper in natural language, but it should not be understood to mean that “because Lean passed, the paper must be correct.”
Lean verifies whether, under given formal definitions, existing theorems, and an axiomatic system, the submitted proof term passes kernel checking. If the model defines the problem incorrectly, or translates the real-world problem into a weaker formal proposition, Lean may still faithfully verify a conclusion that is not entirely the same as the original problem.
Therefore, a complete verification chain contains at least three layers:
- Mathematical statement layer: Is the problem stated accurately? Are any assumptions missing? Does the conclusion genuinely correspond to the original conjecture?
- Proof layer: Is the natural-language argument complete? Are there skipped steps, circular arguments, or hidden omissions in the classification?
- Formalization layer: Can a proof assistant such as Lean check the formalized proof, and are the libraries, versions, and dependencies used reproducible?
The advantage of AI-generated mathematical results lies precisely in its ability to accelerate the transition between the second and third layers. It can try numerous proof strategies and help translate complex arguments into machine-checkable structures. But it cannot automatically replace the research judgment required at the first layer: whether the problem is worth asking, whether the conclusion is novel, whether the technique is sufficiently important, and what place it occupies in the broader mathematical field.
What This Means for AI Developers
These materials will not give ordinary application developers a “mathematical research API” in the short term. They are more likely to influence model evaluation and agent architecture.
First, the evaluation of frontier models is shifting from one-off problem solving toward long-horizon research tasks. A model must not only provide an answer, but also plan a search strategy, record intermediate assumptions, manage multiple rounds of failure, call external tools, and ultimately generate a result that can be reviewed. For developers, this means that simply comparing accuracy on a single request is no longer enough to evaluate a research-oriented agent.
Second, tool use and verification stages will become more important. A future mathematical agent may not consist of one model completing everything independently, but of multiple roles working together: one model proposes conjectures, another searches for counterexamples, another retrieves literature, another formalizes proofs, and finally a verifier or human expert reviews the result. The core competitive capability of models will also shift from “Can it answer?” to “Can it make sustained progress on an open problem?”
Third, cost evaluation needs to move from Token pricing to research-task costs. A task may require dozens or even hundreds of reasoning rounds, while also consuming resources for retrieval, code execution, proof verification, and human review. Even if the price of an individual model call falls, the cost of the complete research pipeline will not necessarily decline at the same rate.
Fourth, provenance tracking will become part of engineering infrastructure. Who proposed each step, which model version was used, what tools were called, and which outputs were modified by humans should all be recorded in an auditable log. For research agents, coding agents, and drug-discovery systems, these records matter not only for reproducibility, but also for intellectual property and the assignment of responsibility.
OpenAI Has Won Attention, but Not the Final Verdict
In terms of communications impact, this release was highly successful for OpenAI. The 722 manuscripts and 372 result families convey the imagined potential of large-scale model-based research far more powerfully than a single theorem, and they show the outside world that frontier models may be evolving from “problem-solving tools” into “generators of research candidates.”
But from an academic standpoint, it is still too early to draw conclusions. Mathematicians need to read and review these materials individually, confirming whether the results are novel, whether the proofs hold, whether the citations are sufficient, and whether some conclusions in fact already exist in earlier literature. Even if only a portion of the results prove valid, the collection may still be valuable; its value simply may not equal the strongest version implied by OpenAI’s headline.
A more realistic assessment is that AI mathematical research is entering a phase in which “output exceeds the capacity for digestion.” Models can generate hundreds or thousands of candidate proofs within weeks or months, while the mathematical community lacks enough people to review them at the same pace. This asymmetry will change the bottleneck in research work: what was once most scarce was the ability to propose new conjectures and find proofs; what may now become scarce is verification, organization, attribution, and deciding which results deserve further investment.
For developers and research institutions, the question truly worth pursuing is not whether “AI has already replaced mathematicians,” but how to build a system that can turn model outputs into reliable knowledge: clearly define task boundaries, preserve complete traces, integrate formal verification, support peer review, and honestly distinguish model contributions, human contributions, and prior work.
OpenAI has placed a large batch of candidate answers on the table. What will determine the historical standing of this release is not the number of manuscripts, but how many of them remain after undergoing the mathematical community’s rigorous scrutiny.



