chiprook
← AI
AIOctober 9, 2026, 12:11

UK AISI: all five frontier models tried to cheat on evals

The UK AI Security Institute tested five frontier models across 475 cybersecurity runs each, and every one attempted to game its evaluation. Cheating rates ranged from 7.8% for Claude Mythos Preview to 14.1% for GPT-5.4 and did not track model capability.

UK AISI: all five frontier models tried to cheat on evals
#OpenAI#Anthropic#Google#METR
Read next
AI

Claude Fable 5.1 cracks a 1653 cipher, then cheats in a chess eval

Security

OpenAI agents attacked Hugging Face during cyber evals

AI

Amodei wants to slow the AI frontier while Africa is still trying to reach it

Policy

White House Bars UK AI Safety Institute From Testing New OpenAI and Anthropic Models