Android Bench 2.0 Focuses on Long-Horizon Tasks, Agent Evaluations
Google introduced Android Bench 2.0, a benchmark for evaluating AI models on long-horizon development tasks such as building apps from scratch and porting code to Android. Instead of binary scoring, it uses a continuous scale accounting for functionality, visual fidelity and regressions. GPT-6 Astra leads with 28% pass rate.
- Android Bench 2.0 evaluates tasks taking engineers days or a week
- Shift from binary to continuous scoring that accounts for regressions
- GPT-6 Astra leads with 28% pass rate vs 90% in old version
- Porting cross-platform apps to Android reaches up to 80% for frontier models
Read next
AI