OpenAI has published the evidence that the rumor mill was waiting for. On October 6, 2026, the company released 722 mathematical manuscripts, organized into 372 families of results, in a public GitHub repository called openai/math, alongside a post titled Sharing AI progress in mathematics. The results come mostly from an unreleased internal frontier model. Earlier today our post on the unconfirmed 400-papers rumor was still asking whether anything would appear; now it has, and the number is larger than the rumor.
This post covers what is actually in the repository, how to read the compute figure OpenAI gives, what Lean certificates do and do not establish, and how to triage a release of this size without drowning in it.

TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is in the release? | 722 manuscripts in 372 families, plus Lean code, reasoning summaries and an overview PDF |
| Who made the results? | Mostly an unreleased OpenAI internal model, per OpenAI |
| How many problems were tried? | About 4,000 |
| What did it cost? | On average, roughly three hours of ChatGPT Pro thinking per result |
| Is everything formally verified? | No. Many have Lean proofs; OpenAI warns some unformalized results could have issues |
| License? | Apache 2.0 |
| Is it peer reviewed? | Not as a body. Nothing in the repository says the full set has been refereed |
| Can I use the model? | No. It is unreleased |
What is in the openai/math repository
According to the repository README, the layout has four parts you will care about:
preprints/: the manuscripts themselves, as PDFs with source files and citation instructions.lean/: a Lean library with the verification configuration, holding the formal proofs for the results that have them.reasoning_traces/: abridged summaries of the model's reasoning for a small number of results. The README lists ten, including the irrationality exponent of π, the Mahler conjectures and Kaplansky's direct-finiteness conjecture in characteristic two. Separately, the README notes that the zeta-function zero-free region and Hodge-conjecture-for-CM-abelian-varieties results were exceptions to the standard procedure, and that the zeta write-up was human edited for readability.overview.pdfandCONTENTS.md: the navigation layer that maps families to manuscripts.
The word "families" matters. 722 manuscripts in 372 families means that on average each family has about two manuscripts, presumably variants, extensions or companion papers on related results. The count of independent mathematical results is therefore closer to the family number than the manuscript number, but the repository, not this article, is the place to confirm how OpenAI groups them. We have not independently verified that grouping.
The README is candid about status. It says results sit at different stages of verification. Some are fully formalized in Lean. Others are not, and OpenAI states plainly that some of those could have issues and that it will try to fix them quickly. That is a better disclosure posture than a bare headline count, and it is the right frame for everything below.
How to read "three hours of ChatGPT Pro"
The most unusual detail in the post is the cost estimate. OpenAI says the model was posed roughly 4,000 problems and that the average result used compute equivalent to about three hours of ChatGPT Pro thinking.
Three cautions make this number useful rather than misleading:
- It is an average over results. It says nothing about the spread. Some results may have taken minutes and some far longer.
- The denominator is attempts. About 4,000 problems were tried and 722 manuscripts came out. Whether a problem counts as solved, partially solved or unsuccessful is defined by OpenAI, so the success rate is not something you can compute from the headline alone.
- It is a translation into consumer units. "ChatGPT Pro thinking hours" is a way to make cost legible to readers who know the product. The model that produced the results is not the model in ChatGPT, so the real hardware cost is not stated.
Still, the figure is the first time OpenAI has framed research output in units a paying user can reason about. It matches a direction we saw earlier in OpenAI's Astra ten-problem release, where cost and compute were already part of the story, and it lines up with the advisory group's call for labs to log cost and compute.
What Lean proves, and what it does not
For readers who do not live in proof assistants: Lean is a programming language and proof checker. If a theorem and its proof compile, every inference step is valid relative to Lean's small trusted core. That removes the worst failure mode of AI-written mathematics, a fluent proof with a hidden gap.
It does not remove all risk. Our checklist for judging a mass release lists the main ones, and they are worth repeating now that real artifacts exist:
| Risk | What to look for in openai/math |
|---|---|
| Wrong statement | A Lean theorem that compiles but states something weaker or different from the actual open problem. Compare the formal statement to the manuscript's own claim and to the original source. |
| Missing formalization | Manuscripts with no Lean file. The README says these are the ones that could have issues. |
| Known results | Results that are new to the model but already in the literature. Check the manuscript's citations. |
| Triviality | Valid but easy results. A count of 722 includes both deep and shallow work. |
| Credit | Whether prior contributions are acknowledged. This has been the main source of dispute since August. |
A reasonable way to triage: start with the families that have complete Lean formalizations, read the formal statement first, then read the manuscript's introduction to see whether the two agree. Leave the unformalized manuscripts for people in the relevant subfield.
How this fits the August and September record
The release is the latest step in a short sequence, and each step has changed what we can check.
August 2026. OpenAI's Astra model produced ten advances with a 249-page manuscript and Lean certificates. We covered what was and was not proved in Astra's ten math proofs.
September 2026. OpenAI claimed more than 100 resolved open problems, including a Navier-Stokes claim, and formed an advisory group at the Institute for Advanced Study. See OpenAI's 100-plus claim and advisory group and the Navier-Stokes agent-swarm announcement. The independent view of that claim is in the explainer on why AI did not simply solve Navier-Stokes.
October 6, 2026. The repository arrives. The README also says OpenAI is exploring community-hosted repositories for these materials and will preserve the public release history, recording corrections as new versions with earlier versions still accessible. We could not access OpenAI's announcement page to confirm further statements about the advisory group or a future model release, so we do not report them.
The group's standards are summarized in AGMAI's rules for releasing AI math: cite prior work, write for humans, deposit in an archive, log cost, formalize, and support exposition. Measured against that list, the repository does several things (public archive, Lean, cost estimate, caveats) and leaves others open (a peer review pipeline, a complete status table per result, and the unreleased model itself).
Why this matters beyond mathematics
It would be easy to treat this as a story for specialists. Two parts of it generalize.
First, it is a model of how to ship an evidence-bearing claim. The release pairs a headline number with a machine-checkable layer and an explicit list of what is not checked. The same pattern is what Anthropic's Lean proof work and the discussion of the collapse in formal proof cost point to: when checking is cheap, claims should arrive with their checkers.
Second, it changes the competitive picture. Google and Meta have both made math announcements this month, covered in Gemini and five math problems and Meta Muse Spark and six open problems. A repository of 722 manuscripts sets a much higher public bar for what a "math result" release looks like, in volume and in disclosure. Labs that announce counts without artifacts will now be compared against this one.
For people who build with or teach AI, the practical lesson is about reading claims. A model that can produce hundreds of plausible manuscripts moves the bottleneck from generation to adjudication, a theme the community has been calling the scarcity of human review. If you evaluate AI outputs in any field, the transferable skills are the ones in the table above: check the statement, find the checker, ask for the denominator, and see who has reviewed it.
What people are asking
Does this prove AI is doing real research mathematics? It is evidence worth taking seriously where Lean certificates exist and the statements match known open problems. It is not evidence about the unformalized manuscripts, and it does not settle questions of novelty or significance, which depend on expert judgment.
Why release now? The timing follows weeks of criticism about announcements without artifacts, plus the rumor that surfaced earlier. The README says OpenAI expanded its open-problem evaluations after existing math evaluations saturated.
Is the model coming? The README does not announce a release of the model. Until one exists the results cannot be reproduced by outsiders, only checked.
What happens to errors? The README says OpenAI will endeavor to fix issues quickly, update the repository as more formalizations arrive, and record corrections as new versions.
How to explore the repository yourself
You do not need to be a mathematician to inspect it.
- Clone or browse openai/math and read
CONTENTS.mdandoverview.pdffirst. - Pick one family with a Lean formalization in
lean/and open the formal statement. - Open the matching manuscript in
preprints/and read the introduction and the citations. - Ask whether the formal theorem says what the manuscript says it says.
- If you find a discrepancy, report it through the repository so the fix is visible to everyone.
Teachers can use the same steps as a classroom exercise in what verification means. Our guide to how language models approach math and science is a gentler background read for non-specialists.
Update (October 7, 2026): the Millennium Prize claims and the repo link
The repository is at github.com/openai/math. OpenAI's announcement post on X, which passed 3.5 million views, says the results were produced by an internal frontier model and that OpenAI consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release them.
Press coverage since then highlights the most ambitious entries. Reports say the release includes claimed progress on three Millennium Prize problems (the Riemann hypothesis, the Hodge conjecture and Birch and Swinnerton-Dyer), a result described as a "quasi-Riemann hypothesis," a proof of NP-hardness for Basic-SDP threshold problems without assuming the Unique Games Conjecture, and the Hodge conjecture for CM abelian varieties, which is a special case, not the full conjecture. Treat those descriptions with care: they come from secondary summaries, "partial progress" is not a solution, and Sam Altman described the results as claims not yet confirmed by outside mathematicians.
One viral post said AI "just solved 90 of the 500 most important open problems in mathematics." We could not find a source for that framing, and nothing in OpenAI's materials says it. Use the checklist in this post, and read the specific manuscript and Lean files before repeating any headline.
What people are saying
Reaction on X has run in three directions, and it is worth separating them from the evidence about the manuscripts themselves. These are public posts, summarized, and none of them is a finding about whether the results are correct.
The empathy argument. Nicolas Bustamante, who works on AI for knowledge workers at Microsoft, wrote that people should "lead with empathy." His point: when you have spent years learning a craft, whether code, illustration or mathematics, and much of it starts to look like "a prompt away," that is disorienting, because the craft is part of how you define yourself. He said he is confident new, harder problems and crafts will appear, but that knowing it intellectually "doesn't make the sense of loss disappear," and that people need time to process it. The post spread widely.
The pushback. Replies were sharply divided. Some called empathy "the new cope." One argued that software engineers have not been reluctant bystanders, since many pushed for AI to take over parts of their own work. Others took a calmer line: good problem-solving tools are a good thing, and the real prize arrives when models at this capability can be inspected and run on affordable hardware, which would let science build on them openly. A sarcastic reply asked whether math is now "obsolete," and another asked whether anyone resents people who calculate with pen and paper.
The extrapolation. One reply wondered what OpenAI might be doing with the same method "for SaaS and jobs." Another viral post claimed AI had "solved 90 of the 500 most important open problems in mathematics." We could not find that framing in OpenAI's materials, and it is a reminder of how quickly a release of this kind turns into a number nobody can source.
Our thoughts: craft, identity and the real bottleneck
We think both camps in that exchange are partly right, and that the more useful question is what changes in mathematics specifically.
Generating results is getting cheap. Knowing which ones matter is not. OpenAI says it posed roughly 4,000 problems and that the average result used about three hours of ChatGPT Pro thinking. If those figures hold, the scarce resource in mathematics shifts from producing candidate proofs to judging them: deciding which statements are interesting, which are new, which are correct, and which are only technically true. That is work for people with deep taste and time, and a flood of manuscripts increases demand for it. The advisory-group process, the Lean certificates and the public repository are attempts to make that adjudication tractable. They will be tested in the coming weeks as mathematicians read.
Formal proof changes what counts as expertise, not whether it counts. A Lean file can confirm that each step is valid. It cannot confirm that the statement is the one the field means, or that a result matters. Reading a formal statement and judging whether it captures the conjecture is a skilled task, and so is explaining why a result is surprising. Fields that adopt machine-checked proof tend to value people who can bridge the formal and the conceptual, much as compilers raised the value of programmers who understand systems.
The identity point is real, and it is not only about mathematics. People who define themselves by a craft feel a threat when a machine does part of it, even if the field grows. Mathematics has an old tradition of valuing understanding over answers: a proof that no human can follow is less satisfying than one that illuminates. One reply put the feeling plainly: they liked it better when unsolved problems were unsolved. That sentiment deserves to be taken seriously by the labs releasing these results, which is part of why the advisory group's call to write for humans, cite prior work and fund exposition matters. A release that arrives without explanation can feel like a loss. A release that comes with readable accounts, credit and invitations to collaborate can feel like a tool.
What we would do about it, practically.
- If you are a mathematician or student, read one manuscript close to your area, in full, with its Lean file. Judge it the way you would judge a colleague's draft. Write down what you would have wanted explained. That judgment is the thing in short supply.
- If you teach, use these papers as case studies in verification: what a certificate establishes, what remains for a human, and how a claim becomes accepted knowledge.
- If you build AI tools, notice that the bottleneck they reveal, adjudication, applies far beyond math. Any system that can produce plausible output at volume needs cheap, trustworthy checking and clear provenance.
- If you cover the story, keep the language honest: "claimed," "partial," "formalized," "unreviewed." The difference between "partial progress on the Riemann hypothesis" and "solved" is the entire story.
- If you feel the disorientation, it is a reasonable reaction. The sustainable response is usually to move toward the parts of the work that depend on taste, judgment and explanation, and to treat the tools as something to understand, not only to fear or celebrate.
None of this resolves whether the specific results are right. That will be answered by the people who read them. What we can say now is that the release makes the human part of mathematics, judging and explaining, more visible, not less.
What is still unconfirmed
To keep this post honest, here is what we could not establish from the primary sources:
- Which of the 722 manuscripts correspond to the 100-plus problems claimed on September 21.
- How many of the 722 are fully formalized in Lean, and how many are not. (By our count, CONTENTS.md links a Lean doc on 235 of the 372 families; our result-by-result check of eight headline results shows the Lean scope can be narrower than the abstract.)
- Whether any independent referee has reviewed the set, or any portion of it, at all.
- The identity and size of the unreleased model, and the true compute behind the three-hours average.
These are answerable from the repository over the coming days as mathematicians read it. We will update this post when the community publishes audits, in the same way it was done for August's ten proofs, which drew a published human audit.
Related reading
- Eight headline results, their Lean status and expert reactions
- OpenAI "400 math papers" rumor: what is verified
- OpenAI says its model solved 100+ open math problems
- OpenAI Astra's 10 math advances: what was actually proved?
- AGMAI's rules for releasing AI-generated mathematics
- No, AI did not just solve Navier-Stokes
- Lean 4 and the collapse in formal proof cost
- Are AI labs hoarding solved math problems?
- Google Gemini and five math problems
Primary sources: OpenAI, Sharing AI progress in mathematics and the openai/math repository.
Details are accurate as of October 6, 2026, the day of release. The repository is being updated, so counts and statuses may change.
