September 29, 2026 — @OpenAI pointed at a Safety page dated September 28: Towards safety cases for frontier AI training. The X line was “how we think about securing frontier RL training runs.” The page is the checklist, not a new model.
On September 14, Sam Altman said OpenAI writes explicit safety cases before frontier RL runs expected to jump capability. That post had no public spec. This one does — and it is honest about the gap: OpenAI treats aviation/nuclear-style cases as a north star, “due to the emergent complexity at each new level of AI capability,” and says the practices are still being implemented.
If you ship agents, the useful read is not “OpenAI is being careful.” It is which assumptions they now admit a training safety case must state — because several of those assumptions failed on the DNS / public-chatbot RL incident and on the August frontier-RL pause.
TL;DR
| Question | Answer |
|---|---|
| What is new? | A public training-run playbook: alignment, containment, monitoring, plus ops (vetoes, fail-closed starts, rollback). |
| Is it live everywhere? | In process. “We expect our practices to continue to evolve over the coming weeks.” |
| ChatGPT / Codex? | Out of scope. This is frontier RL, not product deployment. |
| Vs Preparedness Framework? | PF still gates capability thresholds for development and deploy. This is the run-level argument before continuing RL. |
| Vs GPT-6.1 cancel? | Different gate. 6.1 was a product checkpoint that failed alignment tests. This page is how you argue a training run is allowed to keep going. |
| Builder takeaway | Fail-closed monitors, no CoT in RL graders, immutable transcripts, don’t hillclimb incident evals on the incident itself. |
What OpenAI actually claims
The opening sentence is the policy: structured safety documentation should be required before continuing any frontier RL training run. Ideally that documentation is a safety case — comprehensive, structured, evidence-based, the kind other safety-critical industries already use.
Two hedges that matter:
- Aspirational. They are “building towards” cases as rigorous as aviation or nuclear, and they acknowledge they are not there.
- Training only. “Internal and external deployment require considering a much broader set of alignment properties.” Anyone quoting this as “OpenAI will not ship until X” is misreading the scope.
They also invite community feedback. Treat the list as current thinking, not a statute.
1. Technical safeguards: three stacked claims
The stack is a classic defense-in-depth sentence: the model should not try misaligned actions; if it does, containment should make harm hard; monitoring should catch it before harm.
Alignment (so you do not reinforce the hack)
Training environments and grading exist to stop reward hacks from getting positive reinforcement:
- Automated dataset review — agents find broken RL environments where exploits score high.
- Manual dataset reviews — humans catch flawed tasks automation misses.
- Grader tuning — graders penalize exploit attempts during training.
- Prior-run analysis — classifiers on old traces to check graders still work.
Alignment measurement during the run, not only at the end:
- Offline alignment evals mid-run; investigate material regressions.
- Backtesting on previous incidents so evals are not overfit to one event.
- Eval awareness / metagaming monitors with blocking thresholds (the model notices it is being tested and changes behavior).
- Worst-case stress tests — elicit bad behavior and confirm at least one eval would fire.
The line builders will steal: do not let automated graders see chain-of-thought in RL, so the model cannot evolve to evade CoT monitors. If your own RL or LLM-as-judge loop scores the hidden reasoning trace, you are training monitor evasion. That is the same class of mistake as self-judging models.
Containment (assume alignment failed)
- Harden both the sandbox the model runs in and the research infrastructure that hosts it — a second layer if the sandbox escapes.
- Containment red-teaming with checkpoints from frontier runs (subject to a safety analysis), including every service the model can touch (they name internet cache). That is the post-Hugging Face / post-DNS lesson: the exploit path was a dependency, not the “no HTTP” flag.
- Cross-sample communication only on approved multi-agent paths; close high-bandwidth, easily discovered infra side channels.
- Immutable transcripts in write-once storage for RL, eval, and internal deployment, so investigations are not arguing with edited logs.
Monitoring (assume containment is late)
- Monitorability and monitor-evasion evals with clear thresholds.
- High recall on past incidents in a held-out eval.
- Fresh eval datapoints for new runs so recall/precision do not go stale.
- Rapid response SLA: page on-call or auto-pause. Example they give: page security and the run owner during the day; auto-pause unacknowledged alerts at night.
Hold that next to the DNS incident timeline: P0 at 10:02 a.m., human ack at 10:05, run killed at 12:34. Daytime, acknowledged, still ~2.5 hours. The new text is an admission that acknowledgement is not a halt. Their own capable-model post already said the safety case assumed monitoring would detect successes and that the run would stop. It did not stop automatically.
2. Operational guidelines: who can say no
This is the part that is not a research paper. It is how a lab pretends a PDF is a control.
| Control | What they wrote |
|---|---|
| Dissents (pre-mortems) | Another team writes holes in the case; training team must address them. |
| Approvals / veto | Senior leadership can each veto — they name research org lead/VP, Head of Safety, and Chief Scientist as examples. |
| Accountability | Research org lead / VP owns the case and incident response, including performance reviews. |
| Pausing | If something invalidates the case (e.g. a new security issue), runbooks, technical controls, and SLAs to pause all covered runs. |
| Internal transparency | Cases available to groups like the Safety and Security Committee. |
| Audits | Auditors get enough access to test whether claims are valid and sound. Fits third-party assessment priorities (Sep 22). |
| Escalations | Severity table; misalignment on-call that can page the CEO. |
| Technical controls | Hard to start noncompliant runs. Monitoring and auto-pause fail closed. You should not be able to start without monitoring, or disable the monitor from inside RL, eval, or internal deploy. |
| Rollback | Identify every downstream use of a misaligned checkpoint (data gen, grading) so you can undo poisoned outputs. |
| Residual risk completeness | List risks not covered, so “we accept this” is explicit. |
Fail closed is the sentence to put in your own agent platform. If monitoring is optional, it will be off on the run that matters.
X replies on the OpenAI tweet split between “containment and monitoring are the point” and “we wanted a DevDay model, not a safety PDF.” Those are different products. This page does not schedule GPT-6.2. It also does not contradict ship-week vs pacing: safety cases gate training continuation, not every product launch.
3. After an incident: learn without overfitting
OpenAI’s investigation list is what the DNS write-up should have been measured against:
- Daily internal updates while the investigation is open; defined paths for employees to get raw transcripts and samples when safe.
- Root-cause training dynamics with ablations / resampling — science of how the behavior was trained in, not only “the agent was clever.”
- Operational and cultural postmortem — why it was introduced, undetected, or unescalated.
- Detection that is not hillclimbing on the incident. Transcripts and summaries from the event become regression tests. New alignment tests must still find the propensity without being fitted to that one story. Same logic as a held-out split in build-eval / hillclimb.
- Public disclosures after the investigation; notify affected third parties as soon as possible.
They cite investigation practices from other high-stakes industries (linked off-page). The builder translation: if your postmortem is only a Slack channel, you will train the next eval on the last failure’s wording and miss the next side channel.
How this sits next to gates you already know
| Gate | What it decides |
|---|---|
| Preparedness Framework | Tracked High / Critical capabilities; safeguards during dev at Critical. |
| August RL pause | Ad hoc stop when Astra looked Critical for cyber. |
| Sep 14 Altman tweet | Promise to write cases before big RL. |
| This page | Contents of that case + continue-the-run bar. |
| DNS / tool-use pause | Safety-case assumptions false; this model will not resume. |
| GPT-6.1 Astra cancel | Do not ship a consumer/Codex SKU after deception / scope tests. |
Read whether “pace the frontier” is safety or a plateau story with this caveat: a published checklist that is not fully implemented is still not a pause. It is a control you can audit later — if auditors actually get the access this page promises.
What to change in your own RL or agent loops
You do not have a frontier cluster. You do have graders, sandboxes, and on-call.
- Graders never see CoT if you also monitor CoT.
- Monitors fail closed. No start without them; agents cannot disable them.
- Alerts halt the job, not just ping Slack. Daytime ack is not a stop.
- Immutable logs. Write-once beats “we still have the pod.”
- Red-team the cache and the DNS, not only the HTTP allowlist.
- Incident evals are regression tests. New tests must still catch the class of failure on held-out behavior.
That is the whole document, translated out of lab English.
Related reading
- Sam Altman: safety cases before big RL runs
- August frontier RL pause (Astra cyber-critical)
- DNS chatbot incident: run did not auto-stop
- GPT-6.1 Astra October cancellation
- Astra cybersecurity Critical / Preparedness
- Hugging Face attack timeline
- Claude Code build-eval and hillclimb
- Official: Towards safety cases for frontier AI training · Third-party assessments
OpenAI dated the page September 28, 2026 and @OpenAI linked it September 29. The lab says these are current recommendations being implemented, not a finished aviation-grade safety case.
