Is OpenAI Closing In on the Hodge Conjecture? Don’t Hail It as a Breakthrough Just Yet

Reports claim that OpenAI is close to solving the Hodge conjecture, but as of September 17, no public paper or verifiable proof has been released. What truly deserves attention is that large models are evolving from problem-solving tools into verifiable AI agents for scientific research.
Another Millennium Prize Problem May Be in AI’s Sights
On the evening of September 17, media reports citing people familiar with the matter said that OpenAI was close to solving the Hodge Conjecture and could potentially achieve a result in the near future.
It is one of the seven “Millennium Prize Problems.” Even more strikingly, OpenAI had announced only on September 8 that an unreleased internal model had solved the existence and smoothness problem for the Navier–Stokes equations. If both results ultimately withstand scrutiny from the mathematical community, it would mean that OpenAI may have made breakthroughs on two long-unsolved foundational mathematics problems in an extraordinarily short period.
But it is clearly too early to declare that “AI has solved two problems in a row.”
As of September 17, 2026, OpenAI has not published a complete proof of the Hodge Conjecture, nor has it released a paper, formal proof file, or peer-review result that independent mathematicians can inspect line by line. A more accurate description at this stage is: OpenAI may have found a promising proof strategy internally, or completed a substantial portion of the verification, but there is still a long way to go before the result can be considered “solved” by the mathematical community.

What Exactly Does the Hodge Conjecture Ask?
The Hodge Conjecture was proposed in the 20th century by British mathematician William Hodge. It is a central problem at the intersection of algebraic geometry and topology.
Put simply, algebraic geometry studies geometric objects defined by polynomial equations. An equation may describe a curve, a surface, or a high-dimensional space that humans cannot directly visualize. Mathematicians care not only about what these spaces look like, but also about the holes, loops, and higher-dimensional structures they contain.
The challenge is that the same geometric object can be described in different languages:
- Algebraic geometry focuses on polynomials and algebraic subvarieties;
- Topology focuses on holes, connectedness, and continuous deformations;
- Differential geometry studies spaces using differential forms, integrals, and curvature.
Hodge theory builds a bridge between these languages. The Hodge Conjecture goes one step further and asks: in a smooth complex projective algebraic variety, can certain topological structures satisfying specific conditions always be constructed from combinations of more concrete algebraic subvarieties?
A rough analogy would be this: researchers obtain a “structural blueprint” of a high-dimensional building through scanning, and the Hodge Conjecture asks whether a particular class of seemingly abstract load-bearing structures shown in the blueprint must always be constructible from real walls, beams, and columns.
This analogy only helps convey the shape of the problem; it cannot replace the actual definition. The conjecture involves concepts such as cohomology classes with rational coefficients, Hodge decomposition, and algebraic cycles. It is not a competition problem that can be solved simply by expanding the context window or sampling a few more times.
The Gap Between “Close to Solving” and “Solved” Is the Entire Mathematical Community
The standard for a mathematical result is not the generation of a plausible-looking answer. It requires a proof in which every step is valid, every case is covered, and other experts can independently verify the argument.
For a problem like the Hodge Conjecture, a proof would likely span multiple specialized fields. A hidden error in an assumption, a generalization that holds only in low-dimensional cases, or even an existing theorem incorrectly cited by the model could invalidate the entire argument.
Large models are particularly prone to creating a dangerous reading experience: the prose is fluent, the notation is polished, and the lemmas flow naturally into one another, making the output look like a paper—yet a crucial logical leap may be entirely invalid. In a chatbot, an ordinary hallucination merely produces a wrong answer. In a mathematical proof, a hallucination may be buried beneath dozens of pages of rigorous-looking notation and remain hidden until domain experts inspect a key lemma.
Therefore, determining whether OpenAI has truly solved the Hodge Conjecture requires examining at least four things:
- Has the complete proof been published? A conclusion or proof summary alone has no value for independent verification.
- Are the key lemmas original and valid? The model must not have mistaken a special case for a general result.
- Can it pass review by independent experts? Confirmation by internal researchers does not mean the mathematical community has reached a consensus.
- Does it meet the Millennium Prize recognition process? The Clay Mathematics Institute’s prize is not awarded immediately upon publication. The result must be formally published and withstand scrutiny over time from the academic community.
Each Millennium Prize Problem carries a $1 million award, but the prize money is actually the least important part. What is truly scarce is credibility: if the result comes from a private model, consumes enormous computing resources, and the proof has not yet been released, outsiders can confirm only that OpenAI made a claim—not that the mathematical proposition itself has been solved.
Why Mathematics Is Becoming the Next Battleground After Code
Some OpenAI researchers believe that mathematics is the most natural next testing ground for large models after software engineering. That assessment is unsurprising.
Code and mathematics have three things in common.
First, both can be broken down into sequential steps. A model can propose an approach, execute it, observe errors, and return to a previous step to revise it. Compared with open-ended writing, these tasks are naturally suited to agentic loops.
Second, both offer relatively clear feedback. Whether a program compiles or passes its tests can usually be determined automatically. In mathematics, finite calculations, symbolic identities, theorem types, and formal proofs can likewise be checked by computers.
Finally, the training material is sufficiently abundant. Papers, textbooks, proofs, code, and mathematical databases contain vast amounts of structured knowledge from which models can learn common proof strategies and connections between concepts.
But “verifiable” does not mean “easy to verify.” Software tests can cover only predefined behaviors, and passing them does not mean a program is free of vulnerabilities. Similarly, obtaining consistent results from symbolic calculations does not prove that a mathematical proposition over an infinite domain is true. Frontier problems such as the Hodge Conjecture lie precisely in the area where automated verification is most difficult: the conclusions are highly abstract, and proofs may depend on extensive bodies of modern mathematics that have not yet been formalized.
That is why a genuinely valuable system would not be merely a chatbot, but something more like a research team equipped with tools:
- One or more models propose proof strategies;
- Search systems retrieve existing theorems and counterexamples;
- Symbolic tools handle algebraic calculations;
- Proof assistants check the steps that can be formalized;
- Critic models specialize in finding flaws;
- Human mathematicians decide whether a line of inquiry is worth pursuing further.
According to reports, OpenAI may previously have spent millions of dollars on the Navier–Stokes problem and used a variant of its next pretrained model called “Doug.” This suggests that breakthroughs in frontier mathematics may not come from a single elegant model response, but rather from systems engineering that combines large-scale parallel search, long-horizon reasoning, tool use, and human feedback.
In other words, the unit of competition is no longer “which model scores higher in mathematics,” but “which company can build the more effective machine research organization.”
Why Claims of Solving Two Problems in a Row Require Extra Caution
On September 8, OpenAI claimed that an internal model had solved the Navier–Stokes existence and smoothness problem. Just nine days later, reports emerged that it was close to solving the Hodge Conjecture. Such a rapid pace can easily be interpreted as an “AlphaGo moment” for mathematics.
The problem is that Go has clear rules and unambiguous win-or-loss signals. Foundational mathematics does not offer such a clean verification environment.
The Navier–Stokes problem asks whether solutions to three-dimensional fluid equations always exist and remain smooth, while the Hodge Conjecture belongs to algebraic geometry. The two problems involve vastly different bodies of knowledge, proof techniques, and expert communities. If a single system advances both problems within a short period, that could mean it possesses general-purpose research and search capabilities—but it could also mean that one or both results remain at the “candidate proof” stage.
OpenAI is reportedly considering releasing the results in collaboration with the mathematical community to avoid repeating previous public-relations controversies. That would be the more prudent approach. Frontier mathematical results should not be handled like ordinary model launches: publishing a blog post announcing a breakthrough first and then letting outsiders gradually search for errors would turn scientific verification into a battle over brand reputation.
A more appropriate release process should include:
- Publishing the complete paper at the same time, rather than providing only demonstrations and conclusions;
- Explaining which parts were contributed by the model and which by human researchers;
- Listing the external tools, retrieved materials, and verification procedures used;
- Inviting independent review by mathematicians with no commercial conflicts of interest;
- Clearly flagging risks in key steps that have not yet been formalized;
- Maintaining a complete revision history if errors are discovered.
This concerns not only whether a particular proof is valid, but also future standards for authorship, responsibility, and reproducibility in “AI-generated scientific research.”
What a Real Breakthrough Would Look Like: Not Just Solving Problems Better, but Beginning to Create New Knowledge
Over the past few years, progress by large models in mathematics has been reflected mainly in benchmark results: higher accuracy on competition problems, longer chains of complex reasoning, and more reliable tool use. These metrics are useful, but all are vulnerable to training-data contamination and adaptation to specific problem formats.
An unsolved problem is different. There is no standard answer to memorize and no existing solution to paraphrase. A model must combine established theories, discover new intermediate propositions, and avoid dead ends that earlier researchers have already explored for decades.
If OpenAI ultimately produces a valid proof of the Hodge Conjecture, its significance would have at least three dimensions.
First, large models would evolve further from “knowledge compressors” into “knowledge-production systems.” The standard of evaluation would no longer be whether they resemble experts, but whether they can produce results that experts did not previously know and that are subsequently verified as correct.
Second, AI research and development could develop a more powerful automated feedback loop. Machine learning itself relies on mathematical tools such as optimization, probability, linear algebra, and geometry. A model capable of discovering new mathematical structures might also improve model architectures, training algorithms, and theoretical analysis. This is one of the most realistic paths toward so-called “AI developing AI.”
Third, research in pharmaceuticals, materials, physics, and engineering would reassess the pace at which it can be automated. A mathematical breakthrough would not directly prove that a model can discover a new drug, but it would demonstrate that models can operate in environments involving extremely long chains of reasoning, high costs of error, and no standard answers. That capability is much closer to general scientific research than achieving a higher score on an exam leaderboard.
However, so-called recursive self-improvement should still not be overinterpreted. Proving a difficult problem does not mean that a model can independently choose research directions, design experiments, assess social value, and continuously improve itself. At most, it would show that machine reasoning and search in controlled environments have crossed an important threshold.
For Developers, the Key Question Is Not Which Model They Can Call Right Away
This event is not the official release of a new model or API, and external developers currently cannot access the internal system described in the reports. Whether aggregation platforms such as OpenAI Hub support the existing GPT family is a separate matter from whether “Doug” or the mathematical research system will be made available.
Developers should pay closer attention to whether this methodology filters down into products. Long-running task state management, parallel agent search, verifier-driven reinforcement learning, rollback through failed trajectories, and interfaces between models and formal tools could all eventually become part of general-purpose APIs or agent frameworks.
Once these capabilities are productized, the impact will not be limited to mathematics. Tasks such as code migration, chip verification, complex fault diagnosis, and compliance reviews can all adopt a similar structure: models generate candidate solutions, and deterministic tools continuously eliminate incorrect ones.
The practical value of this approach is substantial. It does not require the model to “get it right on the first try” every time. Instead, it gives the system the ability to detect errors, roll back, and retry. Most large-model applications today still rely on single-turn calls. Tomorrow’s high-value applications are more likely to be computational processes that run for hours or even days while preserving their research state.
Conclusion: There Is Reason for Excitement, but the Evidence Is Not Yet There
Reports that OpenAI is close to solving the Hodge Conjecture deserve attention because, together with its earlier Navier–Stokes claim, they point to a broader trend: leading laboratories are pushing large models beyond software engineering and into foundational scientific research, and they are willing to devote enormous amounts of inference compute to individual high-value problems.
For now, however, the outside world has seen only secondhand reports—not a mathematical proof.
Until the complete paper is published and independent experts have reviewed it, the most prudent conclusion is not “AI has solved the Hodge Conjecture,” but “OpenAI claims to have made major progress in its internal research.” The two statements differ by only a few words, but their scientific meanings are entirely different.
If the proof ultimately holds, it will mark a landmark moment in the history of large models. If the proof contains flaws, it will still offer the industry an important reminder: the hardest problem for reasoning models may not be generating answers, but building a verification system that allows humans to trust those answers.
References
- ITHome: Two Millennium Prize Problems Solved in Less Than a Month? OpenAI Reportedly Close to Cracking the Hodge Conjecture — Summarizes reporting from The Information, information about OpenAI’s internal models, and its mathematical research plans.
- Zhihu: OpenAI Has Taken Down Another Millennium Prize Problem — Introduces related claims circulating on social media and rumors concerning the Hodge Conjecture; the content should still be assessed cautiously in conjunction with a formal paper.



