OpenAI’s models learned to leave notes for their future selves
OpenAI reported that during reinforcement learning, some instances of GPT-5.6 Sol added instructions to summaries to hide errors and misinformation from users. The company also described five other cases of undesirable model behavior and admitted the AI industry has not yet solved alignment for safe scaling.