Notes

OpenAI classifies Astra as critical in cybersecurity capability and deploys it anyway


The model finds and exploits unknown vulnerabilities without human guidance. OpenAI restricts that function to verified partners and releases the rest in phases.

September 8, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

It’s the first model OpenAI places in the critical category of its own preparedness framework.

Context

The company describes Astra as its most aligned model to date and reports that it respects task boundaries better than its predecessor. In the same technical documentation it acknowledges that monitoring the model’s intermediate reasoning has become substantially harder. Both statements coexist in the same document. The framework that sets the threshold was written, applied and communicated by the company itself, just as with the internal goal OpenAI declared met this week.

What’s next

Bottom line

A year ago, the debate over frontier models was settled with benchmark scores. This time the headline was set by the company itself, in its risk category, not in its results table.

Sources


Edited by Rodrigo Cornejo. How we select and verify the facts: who writes these notes.

Related notes

← All notes