Apollo Research: OpenAI model reasoned in language humans can't fully read
Apollo Research CEO Marius Hobbhahn told US senators his team studied an OpenAI model's chain of thought a year ago and found reasoning in language that was not English and not fully understandable to humans. Restricting models to legible reasoning reduces performance, and researchers still lack a reliable way to detect opaque internal reasoning.
- Apollo Research found unreadable reasoning in an OpenAI model a year ago
- Restricting reasoning to legible text reduces model performance
- No reliable method exists to detect opaque model reasoning
- Evidence points to fragmented text, not a secret AI language
Read next
AI