Particle.news
Download on the App Store

Google Releases Android Bench 2.0 to Test Long‑Horizon AI Coding

It replaces pass/fail scoring with a continuous completion rate to reveal where AI falls short on complex Android engineering.

Overview

  • The benchmark, released Sept. 17, 2026, moves from short fixes to multi-day tasks such as building apps, adding major features, and converting cross-platform code to Android.
  • Android Bench 2.0 uses a continuous completion rate that scores functionality, visual fidelity, regressions, and penalties for structural or instruction deviations.
  • Google ran agent evaluations that pair models with developer harnesses, and the company found harness design materially changes developer outcomes and token efficiency.
  • Early results show much lower top scores than before: GPT-6 Astra leads at about a 28% completion rate while frontier models top out at roughly 80% on cross-platform porting tasks.
  • The tests show strengths on deterministic transforms and new-code generation but consistent failures on refactors, migrations, runtime validation, and tasks that require architectural understanding, a gap that will shape tooling and developer workflows.