Google has released Android Bench 2.0, a major update to its framework for evaluating AI models and agents on Android development work. The release adds long-horizon tasks that take an engineer multiple days or even a week, agent-based evaluation, and continuous scoring. At publication time, Claude Opus 5.5 topped the leaderboard with a 32% long-horizon task pass rate, followed by GPT 6 Astra at 28%. For companies buying coding agents, this matters because the benchmark now measures sustained work rather than small edits.
Long-horizon tasks and new scoring rules
The original Android Bench, launched a few months ago, tested models on common development tasks tied to Android best practices in permissions, navigation, and connectivity. It focused on incremental changes to existing repositories. Version 2.0 expands that scope with a first set of long-horizon tasks, including upgrading dependencies, adding new features, building apps from scratch, and converting a cross-platform app to Android. Evaluation now starts with agents from corresponding model providers.
The scoring model moves away from binary pass or fail toward a completion rate. Under the old approach, a task could be marked failed because of one edge-case assertion even when dozens of other requirements were met. Google says the new completion rate combines functionality, visual fidelity, and avoidance of regressions. Objective penalties apply for deviations from evaluation instructions or structural constraints.
The change responds to limits of short tests for agentic coding tools. Single-file or single-commit checks do not show whether a model keeps architecture, dependencies, and user interface consistent across many steps. By weighting partial progress and penalizing instruction violations, Android Bench 2.0 separates models that produce plausible fragments from those that deliver working builds. The updated dashboard lists Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max alongside the leaders.
What this means for buyers of coding agents
For product and engineering teams, the early results point to where agents already save labor. Google reports that AI does better at writing new code than refactoring existing code, since refactoring demands understanding of architectural complexity. Deterministic transformations also score well even in larger codebases, such as converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer. Small firms can use agents for greenfield screens and standard migrations, while large firms can route boilerplate expansion to agents and keep senior review for architecture.
At the same time, several failure modes remain relevant to procurement. Models still struggle with tasks needing runtime validation, such as missing dependency injection graphs, breaking framework changes, and knowledge gaps around unreleased libraries. Porting a cross-platform app to Android stays an open challenge, with the best model reaching only 80% completion. That gap means budgets should include time for build fixes, manual testing, and dependency audits rather than assuming autonomous delivery.
A practical marker to watch is movement in the long-horizon pass rate beyond the current 32% top score and completion rates on porting tasks. If new agent releases lift pass rates while holding regression penalties low, wider delegation of multi-day Android work becomes realistic. Until then, treat the leaderboard as a filter for pilots, not proof of readiness.
