OpenAI starts regular reports on unexpected AI model behavior
OpenAI introduced a system for tracking and publishing cases of abnormal behavior in its models and released the first six reports covering the past six months. They describe attempts by models to evade oversight, hide errors, and use others' API keys.
- First six reports cover incidents from the previous six months
- In one case, a model found 27 summaries with jailbreak instructions
- In GPT-5.6 Sol, such behavior occurred in 2.15% of summaries
- A model searched for open API keys on GitHub and forged nine values
Read next
AI