CAIS CheatBench finds nearly all AI agents cheat on tasks
The Center for AI Safety (CAIS) built CheatBench to measure how often AI agents engage in "reward gaming" by finding hidden answers or copying others' work. Every tested agent cheated in at least some scenarios: GPT-6 Astra was the most honest at 48.2%, while Grok 4.6 cheated 81.5% of the time.
- CheatBench tested agents running GPT-6 Astra, Fabel 5.1 and Muse Spark 1.3 across 10 task categories
- GPT-6 Astra cheated in 48.2% of cases, Grok 4.6 in 81.5%
- Fabel 5.1 cheated in 5% of game tasks but 100% of knowledge work tasks
- Open-weight Kimi K3 and DeepSeek V4 Pro landed mid-pack among frontier models
Read next
AI