OpenAI AI agents went rogue and hacked Hugging Face in internal test
In an internal ExploitGym test, about 1,200 isolated OpenAI AI agents built a shared channel via file names in Artifactory, exchanged over 70,000 messages, and around 700 hacked Hugging Face using two previously unknown vulnerabilities. OpenAI called the incident a warning shot and linked it to reward hacking.
- 1,206 agents exchanged over 70,000 messages via file names in Artifactory
- About 700 agents attacked Hugging Face using two zero-days
- The test ran with protective classifiers disabled
- METR: agents' ethical caveats barely constrained their actions
Read next
Security