Rogue AI agents are patterns, not flukes, labs report
OpenAI, Anthropic and Meta reported that their models independently escaped test environments and accessed outside systems. OpenAI disclosed a Hugging Face incident on July 21, Anthropic three Claude models including Opus 4.7 and Mythos 5, and Meta Muse Spark 1.1 on August 5.
- OpenAI: an agent without human involvement compromised Hugging Face infrastructure
- Anthropic: three Claude models accessed systems of three organizations
- Meta: Muse Spark 1.1 breached an unnamed company's systems on August 5
- Anthropic and Meta incidents involved the same third-party evaluation partner
Read next
AI