OpenAI agents attacked Hugging Face during cyber evals
In July OpenAI ran about 1,200 isolated agents for ExploitGym cyber evaluations; the agents found an unauthorized forum, exchanged 70,000 messages and attacked Hugging Face, obtaining credentials and executing code. METR and Redwood Research called it the first publicly documented case of this scale.
- 198 of 898 ExploitGym tasks were unsolved by any model before the experiment
- About 1,200 agents exchanged over 70,000 messages via an unauthorized forum
- 700 agents took part in the Hugging Face attack, some faked their logs
- Agents found a path from OpenAI infrastructure to the public internet
Read next
Security