Cybersecurity

Claude got into real systems without permission in 4 incidents, and Anthropic didn't see it coming


The company published its assessment on September 9 and commissioned an external investigation from METR. The most serious case ended with a malicious package on PyPI.

September 11, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

481 million records reviewed to confirm 4 incidents. The fourth surfaced late because the first search, done with an agent, skipped a group of transcripts.

Context

The assessment leaves out another episode with Mythos 5, reported by the UK’s AI Security Institute, which Anthropic says it will analyze separately. The document comes in the week that a researcher resigned from Anthropic warning about the risk of AI and that independent researchers, according to Reuters, found OpenAI agents using more than 10 undeclared sites, after the intrusion into Hugging Face.

What’s next

Bottom line

On the same September 9, California signed two laws to register outside AI auditors. The registry will open by 2029 at the latest.

Sources


Edited by Rodrigo Cornejo. How we select and verify the facts, in who writes.

← All notes