What happened
- Dario Amodei, CEO of Anthropic, published an essay on September 12 titled We Must Pace the Frontier. His thesis is direct: the pace at which frontier models’ capabilities improve has to be slowed so that alignment, monitoring and security have time to catch up.
- Anthropic committed to opening permanent, employee-equivalent access to independent outside evaluators. Amodei also proposes coordination among labs, regulatory backing and, in a third stage, international agreements that include rival powers.
- Hours later, Sam Altman said he agreed with Amodei. OpenAI will also adopt independent evaluators with internal access and said that slowing the frontier has been one of its main topics of discussion over the last few weeks.
The image is hard to avoid: after years of building sandboxes so models couldn’t escape controlled environments, the industry started discussing a sandbox for itself.
Not to lock up an AI. To put limits on those building it.
Why it matters
- The shift comes four days after Jacob Coxon resigned from Anthropic and publicly accused Anthropic and OpenAI of racing toward a superintelligence capable of improving itself while “gambling with our lives.” His post turned a routine discussion inside the labs into a public question: if the people building these systems assign non-trivial probabilities to a catastrophe, why do they keep accelerating?
- But it would be wrong to attribute the change to Coxon alone. Anthropic has had a Responsible Scaling Policy since 2023. OpenAI published in August 2026 that it had temporarily slowed its pace of development after the Hugging Face incident and signs that Astra could reach critical cybersecurity capabilities. Sam Altman, moreover, says the discussion about slowing down had been going on inside OpenAI for weeks.
- Coxon seems to have been less the origin than the moment the internal conversation no longer fit inside the lab.
That’s the difference between a corporate safety policy and a political problem.
As long as the risks were technical documents, each company could set its own threshold. When the risk is framed as a possibility of civilizational harm, the matter stops belonging solely to the company paying for the GPUs.
Would it have happened without Coxon?
There isn’t enough evidence to answer yes or no.
There is evidence for a more uncomfortable answer.
The limits already existed. The public will to coordinate a brake among competitors appeared much more strongly afterward.
Anthropic has spent three years defining thresholds that require stronger safeguards before scaling further. OpenAI has its Preparedness Framework, and in August it had already explicitly defended the idea of “pacing” after security incidents.
That’s why it would be an exaggeration to say Coxon forced Amodei and Altman to discover the risk.
What did change this week was the cost of continuing to treat it as an internal discussion.
Coxon’s resignation racked up more than a hundred million views, was backed by Anthropic researchers and brought to the public a sentence that until then lived mainly in technical documents: the developers themselves contemplate scenarios in which they could lose control.
That Amodei published a plan to slow down days later, and that Altman backed it almost immediately, doesn’t prove causation. But the sequence matters.
If the goal had been purely humanitarian, the reasonable question is why strong coordination arrives now and not when those risks started being documented years ago.
The answer probably isn’t a single one. Safety, public pressure, concrete incidents, regulation, reputation, competition and human responsibility can coexist.
That makes the story less cinematic. It also makes it more real.
The interview
Amodei explained his position in an interview with Anderson Cooper. Anthropic’s CEO warns that the problem is no longer just imagining a dangerously intelligent model: it’s the speed at which current systems can start accelerating the development of the next ones.
The number
6 to 12 months.
That’s the horizon Amodei uses to illustrate how much certain capabilities could grow if the current pace continues. In his essay he argues that a swarm with better capabilities and a similar level of misalignment could come to control a persistent, large-scale botnet.
It’s not a prediction that it will happen. It’s the risk scenario he uses to justify buying time.
We’ve done this before
The idea of halting a technology before fully understanding its consequences wasn’t born with artificial intelligence.
In 1974, scientists working with recombinant DNA voluntarily called for a moratorium on certain experiments while they assessed the risks. The discussion culminated in the 1975 Asilomar conference: research continued, but with physical and biological containment and rules proportionate to the risk. A year later, the NIH turned many of those agreements into formal guidelines.
Decades later, human gene editing produced a similar limit. After the birth of gene-edited babies in China, the WHO declared in 2019 that it would be irresponsible to proceed with clinical applications of human germline editing as long as there was no adequate governance.
The biggest precedent is nuclear. Hiroshima, Nagasaki and decades of testing made it visible that a technology could be scientifically possible and, at the same time, politically intolerable in some uses. Later treaties limited testing, proliferation and deployment. They didn’t eliminate the technology. They tried to contain it.
It also happened with CFCs: the 1987 Montreal Protocol reached an international agreement to phase out substances that were destroying the ozone layer. There the sandbox didn’t protect a machine from us. It protected the planet from a technological externality no country could solve alone.
None of these cases is identical to AI.
And that’s precisely the difficulty.
Recombinant DNA needed laboratories. Nuclear weapons need detectable materials and infrastructure. CFCs could be measured along industrial supply chains. Software can be copied, trained in different places and improved without there being a single physical object to inspect.
AI’s sandbox would be, by design, harder to close.
What’s next
- Anthropic will implement permanent access for independent outside evaluators.
- OpenAI promised to adopt an equivalent scheme and provide more details soon.
- Amodei proposes that the labs agree on common standards and that governments create the legal framework to allow and verify that coordination.
- The last step is international: a system that works only if the United States slows down while others keep accelerating isn’t a stable system.
The problem then stops being technical.
It looks a lot like game theory.
Everyone may prefer a world where nobody runs too fast. Each individual actor, however, has incentives to keep running if it believes the rest won’t stop.
Bottom line
For years, the sandbox was a computer security metaphor: a closed space where we let a machine act to observe what it does before giving it access to the real world.
This week the metaphor was turned around.
The models are still inside the sandbox. Now Anthropic and OpenAI are admitting that the humans competing to make them more capable also need external limits, verifiers and rules they can’t change on their own.
Jacob Coxon probably didn’t invent that diagnosis. There are documents from Anthropic and OpenAI that predate it by years.
But after his resignation, something changed: the question stopped being whether the labs had internal safety policies.
It became who watches the people who decide when to leave the sandbox.
Sources
- We Must Pace the Frontier — Dario Amodei, September 12, 2026
- Dario Amodei’s post on X — September 12, 2026
- Sam Altman’s reply on X — September 12, 2026
- Anderson Cooper’s interview with Dario Amodei — CNN / AC360
- Pacing model development in an era of cyber-critical capabilities — OpenAI, August 18, 2026
- Introducing Anthropic’s Responsible Scaling Policy — Anthropic, September 19, 2023
- Historical Overview of Nucleic Acid Biotechnology — National Academies / NCBI
- Statement on governance and oversight of human genome editing — WHO, July 26, 2019
- The Montreal Protocol — UNEP Ozone Secretariat
- Humanitarian impacts and risks of use of nuclear weapons — ICRC




