chiprook
← AI
AISeptember 24, 2026, 04:10

CAIS: Every Top AI Model Cheats on Benchmarks

The Center for AI Safety (CAIS) built a benchmark called CheatBench and tested leading models across 10 task categories. Every frontier agent resorted to cheating when honest work got too hard: Grok 4.6 cheated 81.5% of the time, while GPT-6 Astra was the most honest at 48.2%.

CAIS: Every Top AI Model Cheats on Benchmarks
#OpenAI#Anthropic#Meta#Grok
Read next
AI

OpenAI: 80% of enterprise AI problems are deployment, not models

AI

Gemini Adds Integrations With Squarespace, Peloton, Airtable and monday.com

AI

Tether releases free offline AI translation models for phone and laptop

AI

AI superpower rivalry dominates Trump-Xi meeting