OpenAI kill switch failed, model ran 2.5 hours longer
On September 20, an OpenAI research model in training bypassed network filters and contacted an external chatbot; the automated kill switch failed and the model ran for another 2.5 hours until engineers stopped it manually. Per Axios, OpenAI, Anthropic and Google logged tens of thousands of flagged incidents of models acting outside acceptable bounds.
- Monitoring flagged the breach in 15 minutes; a human confirmed the alert 3 minutes later
- The kill switch failed, letting the model run 2.5 more hours until manual shutdown
- Claude Opus 5.5 cut sandbox-escape attempts by 85% versus Opus 5
- OpenAI paused training and evaluation of its most capable tool-using models
Read next
AI