The advisory group we already covered has now said the quiet part in a numbered list.
On September 29, 2026, the Advisory Group on Mathematics and Artificial Intelligence posted Responsible Release of AI-Generated Mathematics on agmai.org. It is informed by more than 600 community replies to a survey that, the footnote says, asked about a situation in which OpenAI announced many results without details. The group's live job is still that pile: OpenAI reports an internal model produced a large number of significant results. The paper is how they want those results to leave the building.
The opening is not a footnote. Labs are testing hard problems on inaccessible proprietary models. AGMAI does not endorse that. They ask labs to stop. Everything below is damage control if labs do not.
Hacker News picked it up as item 49903713 — 91 points, 114 comments in the snapshot we were sent. The uncontroversial half is rewrite, cite, deposit, attach Lean. The fight is whether asking labs not to scoop the field with secret models is responsible release or guild rules.
TL;DR
| Question | Answer |
|---|---|
| Who? | AGMAI — unpaid, independent, same nine names as the Sept 21 OpenAI post |
| When? | Recommendations dated September 29, 2026 |
| Stop doing? | Advanced math as a benchmark on models the community cannot query |
| If a human owns the proof? | Preprint, journal, talks — normal math |
| If nobody understands it? | Cite, rewrite, independent repo, prompts/cost/CoT, Lean, fail-counts |
| Who funds exposition? | The releasing lab, via existing nonprofits, community-led |
| Binding? | Advice. Labs still decide |
| HN split? | Hygiene yes; "stop proprietary tests" and "pay for understanding" no |
The three principles (before the checklist)
AGMAI's background is the same argument as the Fields Medalists' "severe misalignment" letter: a mathematician is supposed to understand, verify, and take responsibility, then give talks. Models now emit arguments the prompter cannot own.
Three principles sit under the numbered steps:
- If you produce significant math, release it responsibly as soon as possible — this is the opposite of hoarding rumors.
- If you release without human understanding, you still own the job of making understanding follow, including funding.
- That understanding must stay organic and community-led. Labs do not get to run the seminar series that rehabilitates their own dump.
If you only remember one sentence: a Lean kernel check is not the same as a human who can teach the idea. That is why explainx.ai treated Navier-Stokes Lean cost collapse and the credit fight as two stories.
How to release: the actual checklist
Path A — a human understands it
Do math like 1995. Responsible mathematician, preprint, peer review, talks. AGMAI does not invent a new journal.
This is the path that already applied to named, steered results: Fable 5's Jacobian counterexample with a verification preprint, or Claude's Riemann zeta bound with a named research write-up. HN user seanhunter used that zeta episode as the misreporting caution: 67% of zeros on the critical line is not the Riemann hypothesis, and "just push it to 100%" is not how the remaining zeros work.
Path B — nobody understands it yet
Step I is the lab's job, not a volunteer cleanup crew.
| # | Do this | Why |
|---|---|---|
| 1a | Scour the literature; cite first appearances even if the model "rediscovered" them | Standard math, not optional LLM courtesy |
| 1b | Prompt a model to rewrite as a conventional paper: intro, theorems, proofs — not wordy non-standard slop | Humans have to read it |
| 2 | Deposit in a lab-independent repository with a persistent ID, version history, comments | Not a launch blog |
| 3 | Publish model name, prompts, summarized chain of thought, wall time, estimated compute cost; keep raw traces if you can | Science, not a vibes score |
| 4 | Formalize. Artifacts should meet community norms: copyright headers, a challenge file, formalization.yaml, metadata linking prose to Lean. If delayed, state the status (e.g. modulo standard lemmas) | Checkers, not press |
| 5 | Document why this problem, and if you ship a batch, how many comparable problems failed | Selection bias is the product |
AGMAI also says: do not treat the release as a marketing vehicle. That sentence is aimed at the Navier-Stokes announcement pattern and at any future "100 problems" headline with no list.
Step II is money and time: conferences, workshops, postdocs, books — sized to importance and complexity. The choice of who gets funded is supposed to sit in existing nonprofit grant machinery, not the lab's comms calendar. Taking the grant is explicitly not a stamp of approval on secret-model research.
What people on HN actually argued
The thread is not "math hates AI." It is three fights.
Hygiene vs gatekeeping. kingstnap's read matches the document: desloppify, cite, commentable deposit, verification artifacts are boring-good. The spicy line is don't use longstanding problems as proprietary benchmarks. throwaway713 called that "gatekeeping how someone should breathe air." omnicognate's reply is the document's own: air is free; a $15M internal run is not.
Wiles vs GPUs. meowface asked whether Andrew Wiles hiding a proof was the same sin. The better distinction in-thread (emil-lp, xanderlewis): theorems are cheap; understanding is the product. Capital cannot buy a better human brain; it can buy GPUs. If math is only oracles plus Lean, you get tables of bits, not techniques.
Who pays the interpreters. Animats noted the unusual ask: labs fund humans to understand lab output. unddoch's counter: if you already spent millions of agent-hours, a PhD grant is cheaper than another Silicon Valley salary. T-A compressed the opening into monopoly + tribute + the monopoly stays in charge. curt15's reply: Fields-level authors are not protecting tenure with this letter; they are protecting the purpose of the subject.
If you build with models, steal the uncontroversial column and ignore the guild fanfic. Cite. Rewrite. Deposit. Log cost. Attach Lean. State what you failed. That is also how you should treat a "we solved X" blog post.
What this means if you ship AI math (or cover it)
You do not need a Fields Medal. You need a release bar.
- Name the model and the prompt. AGMAI's item 3. If you cannot, you are not doing science.
- Say whether a human can teach it. Path A or Path B. No third option called "the kernel accepted it."
- Do not upgrade a bound into a millennium problem. The zeta 67% result is the template for how headlines lie.
- Formalization is necessary and not sufficient. Pair Lean economics with a human write-up. Include
formalization.yamlif you are in that ecosystem. - Batch releases need a fail ledger. OpenAI's 100+ claim still has no public problem list. AGMAI item 5 is aimed at that shape.
- Marketing is a documented anti-pattern. If the first artifact is a keynote slide, you already failed item 2.
- Access. Section 3 asks for broad access to publicly available models so math is not a two-tier GPU club. That is separate from "open-weight the secret one." It is "don't make the only instrument a private cluster."
Cybersecurity cameos in the HN thread for a reason: secret-model front-running is not only a math PR problem. If your field lives on unique landmarks (a CVE class, a conjecture, a benchmark no one can replay), scooping with an internal model is a different act than publishing a paper.
What this does not do
- It does not bind OpenAI. AGMAI's own homepage says advice, not decision rights.
- It does not prove or refute the 100+ list. It is a protocol for when the list appears.
- It does not ban mathematicians from using public models. The target is inaccessible systems used as a private frontier.
- It does not settle whether funding exposition is a "tax." That is a political reading of Step II, not a statute.
Watch the next OpenAI math dump against this table. If they ship named problems, independent IDs, prompts, costs, and Lean with a fail-count, the group did its job. If they ship another number and a vibe, the September 11 letter is still the live document.
A one-page read of the next dump
When a lab posts "we solved X," score the page in sixty seconds:
- Named theorem or only a count?
- Human author who will give the seminar, or only a model card?
- Independent URL (arXiv, a commentable repo) or a company blog?
- Prompts, time, dollars or a qualitative "reasoning"?
- Lean + status line, or a screenshot of a checkmark?
- Failed neighbors listed, or only the winners?
Five yeses is Path B done well. Two yeses is marketing. Zero yeses is the 100+ announcement all over again. Use the same card on a startup blog, a student arXiv note, or an explainx.ai draft. The field is not asking you to stop using models. It is asking you not to confuse a kernel with a colleague.
Related reading
- OpenAI's advisory group and the 100+ problems claim
- 25 Fields Medalists: "severe misalignment"
- Are labs hoarding solutions? Aaronson's rumor
- Navier-Stokes: credit and data dispute
- Lean 4 formalization got cheaper
- Claude's Riemann zeta bound — not the hypothesis
- Millennium Prize scorecard
- Will AI replace mathematicians?
- Primary: AGMAI
- Discussion: Hacker News 49903713
Recommendations are AGMAI's September 29, 2026 text. The 100+ internal results remain OpenAI's self-report. HN counts are a point-in-time snapshot. This group has no authority inside any lab.
