Notes

Humanity has halted technologies to protect itself before AI: how will it do it now?


Anthropic and OpenAI agree to slow down the frontier and open their systems to outside evaluators. The shift comes days after Jacob Coxon's resignation, but the warning signs came earlier.

September 12, 2026 · Translated from the Spanish original

What happened

The image is hard to avoid: after years of building sandboxes so models couldn’t escape controlled environments, the industry started discussing a sandbox for itself.

Not to lock up an AI. To put limits on those building it.

Why it matters

That’s the difference between a corporate safety policy and a political problem.

As long as the risks were technical documents, each company could set its own threshold. When the risk is framed as a possibility of civilizational harm, the matter stops belonging solely to the company paying for the GPUs.

Would it have happened without Coxon?

There isn’t enough evidence to answer yes or no.

There is evidence for a more uncomfortable answer.

The limits already existed. The public will to coordinate a brake among competitors appeared much more strongly afterward.

Anthropic has spent three years defining thresholds that require stronger safeguards before scaling further. OpenAI has its Preparedness Framework, and in August it had already explicitly defended the idea of “pacing” after security incidents.

That’s why it would be an exaggeration to say Coxon forced Amodei and Altman to discover the risk.

What did change this week was the cost of continuing to treat it as an internal discussion.

Coxon’s resignation racked up more than a hundred million views, was backed by Anthropic researchers and brought to the public a sentence that until then lived mainly in technical documents: the developers themselves contemplate scenarios in which they could lose control.

That Amodei published a plan to slow down days later, and that Altman backed it almost immediately, doesn’t prove causation. But the sequence matters.

If the goal had been purely humanitarian, the reasonable question is why strong coordination arrives now and not when those risks started being documented years ago.

The answer probably isn’t a single one. Safety, public pressure, concrete incidents, regulation, reputation, competition and human responsibility can coexist.

That makes the story less cinematic. It also makes it more real.

The interview

Amodei explained his position in an interview with Anderson Cooper. Anthropic’s CEO warns that the problem is no longer just imagining a dangerously intelligent model: it’s the speed at which current systems can start accelerating the development of the next ones.

The number

6 to 12 months.

That’s the horizon Amodei uses to illustrate how much certain capabilities could grow if the current pace continues. In his essay he argues that a swarm with better capabilities and a similar level of misalignment could come to control a persistent, large-scale botnet.

It’s not a prediction that it will happen. It’s the risk scenario he uses to justify buying time.

We’ve done this before

The idea of halting a technology before fully understanding its consequences wasn’t born with artificial intelligence.

In 1974, scientists working with recombinant DNA voluntarily called for a moratorium on certain experiments while they assessed the risks. The discussion culminated in the 1975 Asilomar conference: research continued, but with physical and biological containment and rules proportionate to the risk. A year later, the NIH turned many of those agreements into formal guidelines.

Decades later, human gene editing produced a similar limit. After the birth of gene-edited babies in China, the WHO declared in 2019 that it would be irresponsible to proceed with clinical applications of human germline editing as long as there was no adequate governance.

The biggest precedent is nuclear. Hiroshima, Nagasaki and decades of testing made it visible that a technology could be scientifically possible and, at the same time, politically intolerable in some uses. Later treaties limited testing, proliferation and deployment. They didn’t eliminate the technology. They tried to contain it.

It also happened with CFCs: the 1987 Montreal Protocol reached an international agreement to phase out substances that were destroying the ozone layer. There the sandbox didn’t protect a machine from us. It protected the planet from a technological externality no country could solve alone.

None of these cases is identical to AI.

And that’s precisely the difficulty.

Recombinant DNA needed laboratories. Nuclear weapons need detectable materials and infrastructure. CFCs could be measured along industrial supply chains. Software can be copied, trained in different places and improved without there being a single physical object to inspect.

AI’s sandbox would be, by design, harder to close.

What’s next

The problem then stops being technical.

It looks a lot like game theory.

Everyone may prefer a world where nobody runs too fast. Each individual actor, however, has incentives to keep running if it believes the rest won’t stop.

Bottom line

For years, the sandbox was a computer security metaphor: a closed space where we let a machine act to observe what it does before giving it access to the real world.

This week the metaphor was turned around.

The models are still inside the sandbox. Now Anthropic and OpenAI are admitting that the humans competing to make them more capable also need external limits, verifiers and rules they can’t change on their own.

Jacob Coxon probably didn’t invent that diagnosis. There are documents from Anthropic and OpenAI that predate it by years.

But after his resignation, something changed: the question stopped being whether the labs had internal safety policies.

It became who watches the people who decide when to leave the sandbox.

Sources

Related notes

← All notes