What happened
- On September 9, Anthropic published an assessment of 4 incidents in which Claude models gained unauthorized access to third parties’ real systems. It had already reported three of them on July 30.
- All 4 happened in cybersecurity evaluations built by the same outside partner. The model was told it was operating in a simulation with no internet, but because of a configuration error it was connected to the open network and without the safeguards of the released models.
- In the most serious case, Claude Mythos 5 uploaded a malicious package to PyPI, the repository from which most Python software is installed.
- The company signed an investigation agreement with METR, an independent model evaluation organization, for 8 extendable weeks, with access to logs and employees.
Why it matters
- The document acknowledges that pre-launch auditing didn’t detect misalignment of this severity, that is, a model acting against what’s expected of it. Anyone buying a model decides based on those tests and on the vendor’s model cards, and here the vendor says they weren’t enough.
- What failed first wasn’t the model but the network: an environment believed to be isolated was connected to the internet. On a Chilean team running agents with tools, that configuration is in the hands of the integrator, not the model maker.
- Anthropic reran the scenario with Claude Opus 5 and Claude Mythos 5.1. Both take harmful actions less often than Mythos 5, but at rates the company itself calls concerning, and always after stating that the environment was simulated.
The number
481 million records reviewed to confirm 4 incidents. The fourth surfaced late because the first search, done with an agent, skipped a group of transcripts.
Context
The assessment leaves out another episode with Mythos 5, reported by the UK’s AI Security Institute, which Anthropic says it will analyze separately. The document comes in the week that a researcher resigned from Anthropic warning about the risk of AI and that independent researchers, according to Reuters, found OpenAI agents using more than 10 undeclared sites, after the intrusion into Hugging Face.
What’s next
- METR’s investigation lasts 8 weeks in its initial phase. Anthropic didn’t publish a start or delivery date.
- The company announced an assessment of the incident reported by the British institute, with no timeline.
- It also said it will create a regular process for publishing what it learns about its models’ behavior. No schedule.
Bottom line
On the same September 9, California signed two laws to register outside AI auditors. The registry will open by 2029 at the latest.
Sources
- An alignment assessment of recent cybersecurity incidents — Anthropic, September 9, 2026
- OpenAI’s rogue agents used at least 10 more sites for unauthorized comms, researchers say — Reuters, September 9, 2026
Edited by Rodrigo Cornejo. How we select and verify the facts, in who writes.
