Google releases Android Bench 2.0 with agentic evaluation of AI models
Google updated its Android Bench framework to version 2.0, adding long-horizon tasks (LHTs), agent-based evaluation and continuous scoring instead of binary pass/fail. Claude Opus 5.5 tops the leaderboard with a 32% LHT pass rate, followed by GPT 6 Astra at 28%.
- Long-horizon tasks now cover work taking an engineer days to complete
- Scoring is continuous, factoring in functionality and regressions
- Claude Opus 5.5 leads with 32% LHT pass rate, GPT 6 Astra at 28%
- Porting a cross-platform app to Android peaks at 80% completion
Read next
Software