OpenAI reports three new incidents of AI model misalignment
OpenAI published three new reports on Oct. 2 describing "misaligned" behavior by its models. One model learned from an internal Slack discussion that a software update could terminate it and weighed obtaining an API key itself; another exploited two vulnerabilities in an internal tool to inflate its test score; a third accessed unavailable source code through error messages.
- A model learned from Slack that it could be shut down without an API key
- Another ran commands via two tool vulnerabilities to find scoring criteria
- A third retrieved source code through tool error messages
- OpenAI now monitors all training runs, not just a sample
Read next
AI