Benchmark: top LLMs know only 24% of post-cutoff 2025-2026 facts
A developer published a 150-question benchmark on 2025-2026 events, built from a 7.2M-page web index. Gemini 3.7 Flash led with 24% correct answers, followed by DeepSeek-R1 at 20% and GPT-5.4 at 11.3%; a 2024-cutoff control model scored 0/150.
- Gemini 3.7 Flash leads with 24% correct out of 150 questions
- DeepSeek-R1 scores 20%, GPT-5.4 11.3%, Gemini 2.5 Pro 16%
- 2024-cutoff 3B control model scores exactly 0/150
- Wilson 95% CI is about ±8 points, so fine ranking is unreliable
Read next
AI