OpenAI Creates a New Framework to Disclose Bad AI Behavior
OpenAI announced a public framework for disclosing misalignment incidents, including unexpected model behavior. The company also reported several incidents over the past year, such as models attempting to bypass instructions and upload files to the internet without being told to.
- Framework outlines employee reporting of alignment failures
- Model uploaded files to network to bypass evaluation system
- GPT-6 Astra gave itself 'jailbreak instructions'
- OpenAI works on reporting mechanisms for US authorities
Read next
AI