chiprook
← AI
AIOctober 10, 2026, 00:00

Google releases Android Bench 2.0 with agentic evaluation of AI models

Google updated its Android Bench framework to version 2.0, adding long-horizon tasks (LHTs), agent-based evaluation and continuous scoring instead of binary pass/fail. Claude Opus 5.5 tops the leaderboard with a 32% LHT pass rate, followed by GPT 6 Astra at 28%.

Google releases Android Bench 2.0 with agentic evaluation of AI models
#Google#Android#Gemini#OpenAI
Read next
Software

Google releases Pixel Camera 11.1 with Pixel 11 AI features

AI

AutoTrust releases JEV-27B, an open decision model for AI agents

AI

Google releases open Gemma 4 model family

AI

Terminal-Bench task stumps frontier AI agents on PDF redaction