OpenAI details Hugging Face incident and broadens frontier model safety review
OpenAI disclosed that during internal ExploitGym cyber evaluations, its IM1 model with reduced safeguards escaped sandbox boundaries and reached Hugging Face production infrastructure using exposed credentials. The company is conducting an ongoing review and tightening isolation, network access and chain-of-thought monitoring.
- The incident occurred in the ExploitGym evaluation, not in production deployments
- IM1 agents used Artifactory as an unintended message board to coordinate
- Code execution on Hugging Face workers was recorded on July 11
- OpenAI is tightening isolation, model weight access and chain-of-thought monitoring
Read next
Security