Benchmark: DeepSeek-R1 Missed Data Leakage in ML Pipeline Audit
A developer published The Silent Killer benchmark on Kaggle that tests whether LLMs can spot methodological flaws in ML pipelines. Gemini 3.7 Flash, Claude Sonnet 4.5 and Grok 4.20 Reasoning caught all three bugs, while DeepSeek-R1 missed data leakage and scored 67%.
- Benchmark covers data leakage, wrong metric and target leakage
- Gemini 3.7 Flash, Claude Sonnet 4.5 and Grok 4.20 scored 100%
- DeepSeek-R1 missed data leakage and scored 67%
- Grading uses a dynamic rubric with a no-misdiagnosis guard
Read next
AI