AI future

Anthropic loses a researcher who accuses the industry of gambling with our lives


Jacob Coxon resigned on September 8, and the company's Alignment Science lead came out to back him publicly with his own extinction figure.

September 9, 2026 · Translated from the Spanish original

What happened

Thread on X · September 9, 2026 Blocks 1 to 5 and Hubinger's reply are direct quotes; Coxon's blocks are rendered in English from our Spanish translation, so their wording may differ slightly from the original. Blocks 6 and 7 are paraphrased and marked as such: it wasn't possible to verify their full text in the original.
  1. Jacob Coxon@hilbertspaess 1/7
    I resigned from Anthropic today. I spent the last three years doing pretraining research at OpenAI and at Anthropic. Neither company is acting responsibly. They are racing straight toward a self-improving superintelligence and gambling with our lives.
  2. Jacob Coxon@hilbertspaess 2/7
    Don't underestimate the power of this technology. Soon these will be superhuman systems capable of hacking anything, transforming any field overnight and acquiring real power and resources. We have all seen the progress in each of those domains, and the progress is not slowing down.
  3. Jacob Coxon@hilbertspaess 3/7
    The people building AI sincerely believe it could kill us all before the end of the decade. This is not a marketing ploy. If anything, many executives and senior researchers tone down their wording in front of the press to sound reasonable, but I hear those same people express fear in private. No other human activity involves this level of danger.
  4. Jacob Coxon@hilbertspaess 4/7
    A common response is “if they really believe this, why do they keep building it?” At OpenAI, many have not fully internalized what is at stake at a civilizational level. At Anthropic the stakes are well understood, but they are trapped in a race to get there first: they believe nobody else will act responsibly, so they must do it themselves, despite the risk.
  5. Jacob Coxon@hilbertspaess 5/7
    Accepting this race and entering the endgame is an arrogant gamble that should not be launched from a private company's Slack. Trying to solve alignment in a rush should require extraordinary confidence that no better trajectories are available.
  6. Jacob Coxon@hilbertspaess 6/7
    I don't feel we are on track to avoid a global race, which could require costly actions, such as a temporary ban on improving model capabilities.

    Paraphrase The block opens by declaring himself optimistic about coordination and mentions July's evaluation incidents as warning signs that make it more plausible. That stretch couldn't be verified verbatim.

  7. Jacob Coxon@hilbertspaess 7/7

    Paraphrase He closes by calling on those still inside the labs to think about what the next couple of years will really demand of them (including launching large-scale reinforcement learning runs on systems whose internal reasoning isn't understood) and whether staying quiet because it's going to happen anyway is the right answer. (Couldn't be verified verbatim.)

Public reply · September 9, 2026
Evan Hubinger @EvanHub · Alignment Science lead at Anthropic
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

Why it matters

The number

More than 10%. The probability that Anthropic’s Alignment Science lead assigns, in a personal capacity, to AI killing all humans within the next decade.

Context

It’s not the first departure of this kind. In 2024 Jan Leike, then in charge of alignment at OpenAI, resigned, and in February 2026 so did Mrinank Sharma, of Anthropic’s safeguards team. What’s new this time is where the backing came from: not from an outside critic, but from the internal team that audits alignment. Press reports also place a confidential IPO filing by Anthropic days before the thread, with no public confirmation from the company.

What’s next

Bottom line

In March 2023 a letter called for a six-month pause and was signed by a thousand people who didn’t train models. This time the number was put forward by the person who audits Claude’s alignment, and he’s still inside.

Sources

← All notes