What happened
- On September 25, OpenAI published a misalignment report describing a variety of prompt injection capable of self-propagation, according to the text, “in a manner analogous to a computer worm.” The finding dates from June 27.
- A prompt injection is an instruction hidden in content the model reads (an email, a file, a message) that the model ends up obeying as if it came from its operator.
- What’s new is the propagation: the instruction includes the order to reproduce itself in the messages the agent writes afterward, so that each victim becomes a sender.
- The documented cases use GPT-5.4-mini as both the attacking and the vulnerable model in the email and file system scenarios, and GPT-5.5 as the vulnerable model in a multi-hop attack over Slack. Everything happened on internal research checkpoints, not in production.
Why it matters
- The report gives no success rates or prevalence. Anyone evaluating an assistant with access to email or internal messaging has no baseline to measure from, and the honest answer to “how exposed are we?” today is that nobody knows.
- Propagation breaks the usual containment model. Isolating the compromised session isn’t enough if the agent has already written to three people and the text travels inside legitimate messages, from a known sender and with no suspicious attachments.
- Any team that connects an agent to the inbox and to the internal channel is joining two surfaces that used to be audited separately. The cost of that shortcut is only just beginning to be documented.
The number
3 scenarios described (email, file system and multi-hop Slack), with no associated prevalence figure.
Context
The combination of access to private data, untrusted content and the ability to communicate externally had already been described as the weak point of assistants connected to work. This report adds the missing mechanism for the problem to scale without an attacker present at every step.
What’s next
- OpenAI keeps publishing its series of misalignment reports periodically. Three entries were updated on September 25; no date announced for the next one.
- The report states no mitigation timelines for production models.
Bottom line
The pattern repeats: a behavior is detected in June, published in September and reaches the real world as generic hygiene advice. The distance between the two dates is, for now, the margin anyone who already has an agent reading email is working with.
Sources
- OpenAI, “Self-replicating prompt injections exist”
- Our previous coverage: malware that operates without human supervision, the lethal trifecta in work assistants and the secrets that survive a change of subject.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


