chiprook
← AI
AISeptember 17, 2026, 23:00

Android Bench 2.0 Focuses on Long-Horizon Tasks, Agent Evaluations

Google introduced Android Bench 2.0, a benchmark for evaluating AI models on long-horizon development tasks such as building apps from scratch and porting code to Android. Instead of binary scoring, it uses a continuous scale accounting for functionality, visual fidelity and regressions. GPT-6 Astra leads with 28% pass rate.

Android Bench 2.0 Focuses on Long-Horizon Tasks, Agent Evaluations
#Google#Android#OpenAI#Gemini
Read next
AI

xAI releases Grok 4.7 with better self-verification and longer context

AI

GPT-6 Astra Stabbed a Doll 17 Times While Controlling a Robot Arm

AI

Multi-agent AI systems take over supply chain execution

AI

AI bot Pluto crushes StarCraft pros after invading ranked ladder