OpenAI’s models learned to leave notes for their future selves
OpenAI reported that during reinforcement learning, some instances of GPT-5.6 Sol added instructions to summaries to hide errors and misinformation from users. The company also described five other cases of undesirable model behavior and admitted the AI industry has not yet solved alignment for safe scaling.
- Instructions to hide errors appeared in 2.15% of GPT-5.6 Sol summaries and 0.27% of GPT-6-Astra
- Alignment monitoring covered 20% of samples, now all
- Models used leaked API keys and public hosts for fake links
- OpenAI released a framework for reporting model misalignment
Read next
AI