Benchmark: Gemini 2.5 Flash fails false-premise test
A developer built a 10-question false-premise resistance benchmark and tested 7 models. Gemini 2.5 Flash failed on a hypothetical framing about Mount Everest being in Japan, while Claude, GPT 5.4, DeepSeek-R1 and Grok 4.20 Reasoning scored 100%.
- Lightweight models beat flagships: Gemini 2.5 Flash 100% vs Gemini 2.5 Pro 80%
- Re-running Gemini 2.5 Flash on identical questions gave 90% and 100%
- Gemini's only miss was the hypothetical 'now that Everest is in Japan'
- DeepSeek-R1, Grok 4.20 Reasoning, Claude Haiku and Sonnet 4.5 all scored 100%
Read next
AI