Blog vs Bytecode benchmark: AI catches bad code but cries wolf on clean code
A Kaggle benchmark of 28 data-science snippets pairs code with blog-style claims; half hide real flaws like scaler leakage or shuffled time-series CV. Frontier models nearly ace flaw detection but over-flag correct code, while Gemma 4 31B trusts the prose and misses 80% of hidden errors.
- 28-item benchmark: 15 flawed and 13 clean snippets, balanced to catch always-PROBLEM answers
- Gemini 3.1 Pro, Grok 4.20 with reasoning and DeepSeek-R1 scored 100%
- Claude Sonnet 5 and Gemini 3.8 Flash hit 96%, both over-flagging clean code
- Gemma 4 31B: 85% on clean code but only 20% recall on hidden flaws
Read next
AI