OpenAI sets principles for independent AI safety reviews
OpenAI outlined how independent organisations should examine the safety of its most advanced AI models, including access to technical safeguards, visible chains of thought and internal systems. Reviews would cover four areas and run for weeks or months. OpenAI has not named the assessors or said when reviews could begin.
- Reviews would cover four areas: safety cases, jailbreak defences, chemical, biological and cybersecurity risks, and serious model misalignment
- Assessors would get access to technical safeguards, visible chains of thought, confidential data and red-team systems
- Reviews would not be tied to individual product launches but could inform pre-deployment decisions
- OpenAI is discussing proposals with third parties but has not named assessors or start dates
Read next
AI