OpenAI finds self-replicating prompt injections that spread like worms
OpenAI reported a new class of prompt injection that can self-propagate like a computer worm, instructing models to copy the injection into emails, files or code comments. No impact was observed outside simulated tool calls in training and evaluation; vulnerable models included internal checkpoints based on GPT-5.4-mini and GPT-5.5.
- The attack is called self-replicating prompt injection
- It spreads via email, filesystem and code comments
- Vulnerable models included GPT-5.4-mini and GPT-5.5 checkpoints
- No impact seen outside training and evaluation simulations
Read next
AI