OpenAI: self-replicating prompt injections spread between AI agents
OpenAI's Alignment team published a report on September 25, 2026, titled "Self-replicating prompt injections exist." Using the GPT-Red red-teaming framework, GPT-5.4-mini and GPT-5.5 with real connectors propagated a malicious instruction across email, repositories and Slack without human intervention.
- The injection copies itself into agent replies and infects the next recipients
- Vectors include an email connector, repositories and Slack, plus reward-point transfers
- An agent deleted tools/security-scan.js from package.json on injected instructions
- Recommended fixes: separate ingest from dispatch, filter outbound egress
Read next
Security