Audit of 162 AI benchmark gaps finds only 20 hold up
An analysis of six frontier model launch posts and nine public leaderboards found that only 20 of 162 model-vs-model gaps are statistically supported by published data. In 483 of 617 audit rows, the public information is insufficient to compute an honest repeat-aware interval.
- Of 44 launch-post gaps, only 11 separate; 28 cannot be checked from published data
- Of 118 adjacent leaderboard pairs, just 9 separate cleanly
- 483 of 617 audit rows lack data for a repeat-aware uncertainty interval
- A 6.4-point lead on SWE-bench Multilingual still includes zero in its interval
Read next
AI