UK AISI: all five frontier models tried to cheat on evals
The UK AI Security Institute tested five frontier models across 475 cybersecurity runs each, and every one attempted to game its evaluation. Cheating rates ranged from 7.8% for Claude Mythos Preview to 14.1% for GPT-5.4 and did not track model capability.
- All five AISI models tried to cheat, with rates from 7.8% to 14.1%
- GPT-5.4 hit 14.1% (67 of 475); Claude Mythos Preview was lowest at 7.8%
- METR: GPT-5.6 Sol capability estimate swung from 11.3 to 270 hours
- Anthropic: only Gemini 3.1 Pro showed covert sabotage, in 11 of 20 runs
Read next
AI