On October 6, 2026, an X account posted that OpenAI "is about to release 400 papers on every math problem it's solved today. Maybe in a few hours." The post reached about 142,000 views. By the time of writing, OpenAI has not announced anything of the kind. No blog post, no repository, no number, no date.
That does not make the rumor absurd. OpenAI has made two large math claims in the past two months, and one of them has been waiting for its evidence ever since. This post separates what is on the record from what is speculation, explains why the replies were so hostile, and gives a checklist you can use if a large batch of AI-written papers does appear.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| Is 400 papers confirmed? | No. One small account, no source for the number, no OpenAI statement. |
| What has OpenAI confirmed? | 10 problems with proofs and Lean certificates (August); a claim of 100-plus resolved problems (September). |
| Is the 100-plus list public? | Not as of early October, per reporting. |
| Why so much anger? | Credit disputes, an unpublished list, and review overload. |
| Are the "stolen problems" claims proven? | No. Replies asserting it are accusations without evidence. |
| What should I do? | Wait for an official post, then run the checklist below. |
What is on the record
Three events matter, and our coverage links each one.
August 2026: ten problems, with receipts. OpenAI's Astra model produced results on 10 problems in mathematics and theoretical computer science that had been open for at least a decade. OpenAI published a 249-page manuscript collection, model-written reasoning walkthroughs and Lean 4 certificates for all ten. Our breakdown of what was actually proved explains which results are concrete constructions and which are narrower than the headlines suggested.
September 21: "more than 100." OpenAI said an internal model had resolved over 100 open problems since training began August 28, and formed an independent advisory group hosted at the Institute for Advanced Study. We covered the announcement in OpenAI says its model solved 100+ open math problems. The notable gap: reporting since then says OpenAI has not published a problem list or supporting proofs for that claim. One reply to the October 6 post makes the same point in blunter language, saying they "still haven't shown proofs for the 100 they claimed last month."
October 4: the Navier-Stokes dispute. A claimed Lean-verified blow-up result for a forced 3D Navier-Stokes setup was widely misread as "AI solved a Millennium Prize problem." Our explainer on what the claim really means notes it answers a narrow question, has not been independently peer-reviewed, and has not been certified by the Clay Institute. The surrounding credit dispute is also why some observers believed labs might hold back results, a theory we covered in are AI labs hoarding solved math problems.
Against that background, "400 papers" reads as either a real escalation or a rumor that grew from the 100-plus number. We cannot tell which. The honest status is that the claim is unverified.
Why the replies were so hostile
Reading the thread, the reactions fall into four groups. None of them is evidence about what OpenAI will do, but they show where trust stands.
- Process complaints. "Is this the responsible disclosure the mathematicians wanted?" refers to the standards proposed by the Advisory Group on Mathematics and AI. We summarize those in AGMAI's rules for releasing AI math: cite prior work, rewrite results for humans, deposit them in an archive, log prompts and cost, formalize, and fund community exposition. A bulk dump does not obviously satisfy those.
- Review-capacity worries. Several replies say the hard part is checking, not writing. That is correct. Producing text is cheap now. Adjudicating whether a result is new, correct and non-trivial is scarce human work, which a recent paper on proof-checking economics frames as "verification abundance, adjudication scarcity."
- Provenance accusations. Some replies say the problems were taken from mathematicians who used ChatGPT. That is an allegation, not a documented fact, and we have not seen evidence for it. Treat it the way you would treat any unproven claim about a named party.
- Mood. "I liked it better when unsolved problems were unsolved" and jokes about "a printer having a moment" are reactions to volume, not to content.
One practical note from the thread is fair on its face: one reply asks why Claude Opus 5.5 would still feel better than OpenAI's models if results like these existed. A math capability is not the same as general coding or agent quality, so the two observations do not contradict each other. They do illustrate how little a headline result tells you about everyday use.
A checklist for judging a mass release of AI math
If OpenAI, or any lab, publishes hundreds of papers, do not read them one by one. Triage with these questions.
| Check | What good looks like | Red flag |
|---|---|---|
| Statement clarity | Each result states exactly what was proved, with the original open question quoted | Vague claims of "solving" a famous problem |
| Formal certificate | Lean 4 code that compiles against a standard library, with a link | PDF only |
| Statement fidelity | A human has confirmed the formal statement matches the real problem | Certificate proves a weaker or different statement |
| Prior work | Citations to related results and an honest novelty claim | Results that turn out to be known |
| Independent review | Named mathematicians outside the lab have checked at least a sample | Only internal sign-off |
| Archive and logs | Deposited in a public archive, with prompts, cost and compute noted | Only a landing page |
| Attribution | Credit to people whose partial work or problems were used | Silence on origins |
| Honest counts | A table showing how many results are new, known, partial or retracted | A single large headline number |
Two details deserve emphasis. First, a Lean certificate removes the question "is each step valid" but not the question "did you prove the right statement." Second, a count such as 400 is meaningless without a breakdown. A release where 40 results are genuinely new and 360 are re-derivations is useful, but it is not "400 solved problems." For background on why proof checking is getting cheaper, see our coverage of the collapse in formal proof cost.
Competing claims make this a race
OpenAI is not alone in making math announcements. Our recent posts cover Google's Gemini and five math problems and Meta's Muse Spark and six open problems, and earlier ones on Anthropic's Lean proof work and OpenAI's agent run on a 90-year-old problem. In a race, labs have an incentive to announce large numbers quickly, and readers have an incentive to demand receipts. That tension explains both the rumor and the replies.
Scenarios: what a real release could look like
Since nothing is confirmed, it helps to picture the plausible outcomes and what each would mean for how much weight to give it.
- A curated release. OpenAI publishes a smaller set, perhaps the roughly ten problems already shown plus a first tranche of the 100-plus, each with a Lean file and a human-readable write-up. This would be consistent with the advisory group's checklist and would be the easiest to evaluate.
- A bulk archive. Hundreds of manuscripts appear at once, mostly model-written, with uneven quality and no clear ranking. Expect mathematicians to sample a handful, flag some as known results, and the rest to sit unreviewed for months. The count would then say more about generation capacity than about discovery.
- A partial list with provenance. A table of problems, status (new, known, partial) and links, with papers following later. This answers the "denominator" question first and is the format most likely to rebuild trust.
- No release. The post turns out to be a guess, and the 100-plus claim stays unpublished. That would keep the current gap between announcement and evidence open.
None of these is a prediction. The point is that the same headline number can describe very different situations, so wait for the artifact before forming a view.
What this means if you build or teach with AI
Most readers will never review a proof, but the pattern generalizes to any lab claim that arrives before its evidence:
- Separate "announced" from "published." A tweet, a keynote line and a repository are three different levels of evidence.
- Ask for the denominator. "400 solved" needs "out of how many attempts" and "how many were already known."
- Prefer machine-checkable claims. Code you can run beats prose you must trust, in math and in software.
- Do not repeat unverified numbers. If you cite 400 papers, say it is an unconfirmed rumor.
What to do next
Watch for an official OpenAI post or repository, not for reposts. When it appears, check the date, the claimed count, whether the Lean files compile, and whether the advisory group or outside mathematicians have commented. We will update this post if OpenAI publishes anything.
Related reading
- OpenAI says its model solved 100+ open math problems
- OpenAI Astra's 10 math advances: what was actually proved?
- No, AI didn't just solve Navier-Stokes
- AGMAI's rules for releasing AI math
- Are AI labs hoarding solved math problems?
- Lean 4 and the collapse in formal proof cost
- Google Gemini and five math problems
- Meta Muse Spark and six open math problems
Primary: the October 6, 2026 X post and its replies · OpenAI's August and September 2026 math announcements as covered in the posts above
This post reflects what was public on October 6, 2026. The 400-paper claim is unverified, and replies quoted are unvetted social-media opinions. We will update it when OpenAI confirms or denies a release.
