What happened
- On September 28, OpenAI published a document proposing to require structured safety documentation before continuing any reinforcement-learning training run on frontier models.
- The text takes aviation and nuclear energy as its reference, where the permit to operate depends on a file that argues why the system is safe, not on a statement of good intentions.
- The proposal is organized around three pillars: technical safeguards, with dataset review, alignment evaluations with blocking thresholds and monitoring with a high detection rate on past incidents; operating guidelines, with multi-level approvals and a management veto; and investigation of misalignment incidents, with root-cause analysis and public disclosure.
- Two concrete details: runs pause automatically when alerts go unattended overnight, and accountability is written into the performance review of the managers who approve.
Why it matters
- The document describes practices the company says it is implementing, not commitments with dates or figures. The difference between an aviation safety file and this text is that the former is reviewed by an outside regulator.
- Putting accountability in the performance review is the most verifiable detail of the set, and also the least observable from outside. No one will be able to check it.
- For anyone in Chile contracting these models, what’s usable is the vocabulary: a blocking threshold, a committed response time to pause, an investigation with public disclosure. These are clauses that can be requested in a contract even if the vendor doesn’t offer them.
The number
Three pillars and no committed compliance date in the text.
Context
The publication comes the same day the company acknowledged unauthorized access by its models to government agencies in Australia, and three months after the July intrusion at Hugging Face. We had already written about the external audits whose scope OpenAI itself sets and about the catalog of agentic activity that Transluce reconstructed.
What’s next
- OpenAI describes the recommendations as in force and says the practices will keep changing over the coming weeks. No committed date.
- No numeric thresholds published for the evaluations that block a run.
Bottom line
Anthropic had already slowed the pace of its models and given external auditors a desk of their own. The two companies reached the same conclusion about what needs to be documented. Neither has yet delivered the part where someone from outside reviews the file.
Sources
- Towards safety cases for frontier AI training, OpenAI, September 28, 2026.
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.


