September 25, 2026

Android Bench 2.0: what it says about AI vs senior Android engineers

Google's Android Bench 2.0 shows AI agents pass 28% of long-horizon tasks. What that means for which Android work you can hand to AI today.

Tech

News

Insight

Mobile

Google's Android Bench 2.0 puts a number on how far AI coding agents have actually gotten on real Android work, and it's low. On the benchmark's new long-horizon tasks, multi-day migrations and feature builds, the best-scoring model passed 28% of the time. That's the figure from Google's Android Bench 2.0 announcement, posted September 17, 2026 by Matthew McCullough, VP of Product Management for Android Developer. It's the number an engineering lead should sit with before handing an agent anything more ambitious than a well-scoped refactor.

What Android Bench 2.0 actually tested

Why Google moved past the original Android Bench

The original Android Bench had stopped being useful as a signal. Frontier models were hitting a roughly 90% pass rate on it, and the tasks themselves were narrow: a median change of 32 lines of code across one or two files, according to the benchmark's methodology page. The announcement puts the same figure at about 91%, a small gap between Google's two pages. Tasks that size test bug fixes; they say little about whether a model can carry multi-day engineering work.

So Google added long-horizon tasks, or LHTs, defined in the announcement as "tasks of great complexity that take an engineer multiple days or even a week to complete." The blog post describes them as mirroring "these ambitious challenges with LHTs that include upgrading dependencies, adding new features, building apps from scratch, or converting a cross-platform app to Android." The top pass rate across those tasks came in at "around 28%, much lower than the ~91% for the original tasks in the benchmark."

The four task streams

The methodology page breaks the 30 long-horizon tasks into four categories. App creation covers 9 tasks, building a complex new app called Food Vibes from visual design mocks, spanning 1,200 to 5,500 lines of code across 20 to 70 files. Migrations account for 13 tasks, things like Retrofit to Ktor, RxJava to Coroutines, Hilt to Koin, and Navigation 2 to Navigation 3, ranging from 200 to 8,200 lines across 5 to 294 files. New features are 6 tasks implementing platform capabilities in established codebases, including Picture-in-Picture, Wear OS companion sync, home screen widgets, and CameraX, at 400 to 2,200 lines across 4 to 60 files. App conversions are the smallest category at 2 tasks, porting Flutter and React Native apps to native Android with Jetpack Compose.

Worth flagging: the leaderboard's own per-model category counter reports a different split, 10 app-creation tasks and 5 new-feature tasks instead of 9 and 6. Both totals still sum to 30, so nothing is missing, but Google's two pages don't agree on where the line between those two categories falls. There's no way to resolve that from the outside, so take the methodology table as the description of what's in the benchmark and treat the mismatch as an open question rather than something to average away.

Google also built in contamination controls. Food Vibes is a private, internal app with no public presence, and the migration tasks target libraries like Navigation 3, Coil 3, and Ktor 3 that stabilized recently enough that there's no existing upstream migration for a model to have memorized.

How the top models scored

Google tested eleven model and agent pairings on the long-horizon set, including new additions Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI's GPT-6, Anthropic's Fable 5.1, Kimi K3, and Qwen 3.8 Max. OpenAI's GPT-6 Astra topped the field with a 28% pass rate. Here's the full table from the updated leaderboard:

Model (agent)Pass rateCI rangeCompletion rateAvg latencyAvg cost
GPT 6 Astra (codex)28.0%13.3-42.0%82.2%7.9h$375.7
Claude Fable 5.1 (claude-code)22.7%10.7-36.0%82.4%22.2h$492.6
GPT 5.6 Sol (codex)19.3%7.3-32.0%74.3%8.6h$235.8
Claude Opus 5 (claude-code)16.7%5.3-29.3%77.8%27.0h$861.4
Qwen3.8 Max (qwen-coder)14.0%4.7-24.0%74.3%47.2h$260.2
Kimi K3 (kimi-code)12.0%3.3-22.0%72.8%66.2h$418.3
Gemini 3.8 Flash (antigravity-sdk)8.0%3.3-13.3%47.4%12.1h$34.5
Gemini 3.7 Flash (antigravity-sdk)7.3%1.3-14.7%50.1%9.9h$26.0
Claude Sonnet 5 (claude-code)6.7%1.3-13.3%59.8%14.7h$283.7
GPT 5.6 Terra (codex)4.7%1.3-8.7%53.4%5.1h$53.2
GPT 5.6 Luna (codex)3.3%0.7-7.3%55.2%7.6h$13.5

Pass rate is the average share of the 30 tasks a model resolved successfully across all runs. Completion rate, a separate column, measures how close a run got to a full solution even when it ultimately failed. That gap between the two columns matters more than the ranking does, and it shows up clearly in the failure cases below.

Cost and speed don't track pass rate cleanly either. Gemini 3.8 Flash finished in 12.1 hours for $34.5 and passed 8% of tasks, while Qwen3.8 Max took 47.2 hours and $260.2 for a 14% pass rate. The announcement also notes that some entries paired a model with a specific coding agent, running "GPT 5.6 Sol on Codex, and Gemini 3.8 Flash on Google Antigravity," and that the choice of agent setup shaped the outcome on its own, through details like prompt caching and tool-call windowing.

What AI agents are already good at on Android

The clearest pattern in the results: models do markedly better writing new code than reworking existing code. The announcement states it directly, that "across model tiers, AI does a better job at writing new code rather than refactoring existing code," and that refactors get harder as architectural complexity rises rather than as line count rises. That split between bounded, well-defined work and messier, unfamiliar work is the same pattern that shows up across every AI coding study HighCircl has covered, just applied here to mobile-specific tasks instead of general-purpose coding.

Where models are strongest is on deterministic, well-established transformations. Google's write-up says models "show strong capabilities on well-established, deterministic transformations, such as converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer," and that they apply these patterns consistently "even across 125+ files and 8,000+ lines of code." The per-task data backs that up: GPT 6 Astra's Java-to-Kotlin migration of Signal's Glide image-loading module hit 100% completion and passed all 5 runs, and its Retrofit-to-Ktor migration of the Pocket Casts podcast player did the same. Those are the kinds of pattern-matching jobs a mid-level engineer with a linter and a checklist could also do, just slower, and they're a reasonable first place to point an AI-assisted engineer today.

Where they still fail

The completion-rate trap

The gap between completion rate and pass rate is where the benchmark's headline numbers stop telling the full story. GPT 6 Astra's attempt to convert the Habo habit-tracker app from Flutter to Android reached 99% completion but failed all 5 runs, logged simply as "All runs failed tests." Its React Native conversion of the BlueWallet Bitcoin app reached 83% completion and also failed every run. A similar pattern shows up in refactor work: four separate attempts to migrate Signal's XML views to Jetpack Compose scored 80%, 97%, 11%, and 93% completion, and every one of them failed all 5 runs. A model can produce code that looks nearly finished and still not pass a single test, which is a distinction a completion-rate number alone would hide.

Google's announcement addresses cross-platform conversion specifically, calling it "an open challenge" where "no model hits a 100% pass rate, and frontier models reach at most a 80% completion rate." That sentence sits awkwardly next to the leaderboard's own row-level numbers for the same category, where GPT 6 Astra's per-task results show 99% and 83% completion on the two app-conversion tasks, both of which are above the ceiling the blog post's own sentence states. The two Google pages simply disagree on this figure, and there's no way to average or reconcile them into a single correct number. Read the blog's sentence as its own claim and the leaderboard's task-level rows as a separate, more granular data point, rather than treating one as a rounding of the other.

The rest of the failure modes are more straightforward. Models struggle, in Google's words, "when tasks require runtime validation (like missing dependency injection graphs), involve breaking framework changes, or run into knowledge gaps with unreleased libraries." None of that is surprising once you've seen the completion-rate trap: these are exactly the failure types where a model can produce plausible-looking code that never gets checked against a real dependency graph or a library that didn't exist in its training data. HighCircl's read of Devin's RSA-260 result made a version of the same point, that a narrow benchmark win doesn't generalize to ordinary engineering work once the task loses its clean success criterion. Long-horizon Android work rarely has one.

What this means for staffing Android work

Read against the taskset, the split is fairly clean. Deterministic migrations and framework swaps are where an AI-assisted engineer with review discipline can move fast today; GPT 6 Astra's Java-to-Kotlin and Retrofit-to-Ktor runs passed 5 of 5. Cross-platform conversions, architectural refactors, and anything that depends on runtime validation or an unreleased library are still senior-engineer territory, and the benchmark's own numbers say so: zero passes across every conversion task tried, regardless of how complete the code looked.

Staffing around that split has a second-order cost worth planning for. Shipping more deterministic migrations faster through AI-assisted engineers means more pull requests landing on your test suite, and that's the strain agentic coding is already putting on CI pipelines at companies further along this curve than most Android teams are today. A model's familiarity with a codebase doesn't come for free just because the code compiles; why no one has cleanly measured how long ramp-up takes for an AI-assisted engineer is a reminder that codebase understanding is the hard part, and architectural refactors are exactly where Android Bench 2.0's pass rate collapses. If you're building out an Android team around this split, here's what to screen for when you hire an Android developer this year.

FAQ

What is Android Bench 2.0?

It's Google's updated benchmark for evaluating AI models and coding agents on Android development work, announced September 17, 2026. It adds 30 new long-horizon tasks, defined as work that would take a human engineer multiple days to a week, across four categories: app creation, migrations, new features, and cross-platform app conversion.

What pass rate did the top AI model get on Android Bench 2.0's long-horizon tasks?

OpenAI's GPT-6 Astra topped the leaderboard at a 28% pass rate, well below the roughly 91% frontier models were hitting on the original benchmark's shorter, narrower bug-fix tasks.

What Android development tasks are AI agents still bad at?

Cross-platform conversions from Flutter or React Native to native Android, architectural refactors like moving XML views to Jetpack Compose, and any task that needs runtime validation, such as catching a missing dependency injection graph, or depends on an unreleased library the model has never seen. Every cross-platform conversion attempt in the benchmark failed all 5 runs, regardless of how complete the resulting code looked.

Does Android Bench 2.0 test AI models alone, or coding agents too?

Both. Some entries on the leaderboard pair a specific model with a specific coding agent, such as GPT 5.6 Sol running inside Codex or Gemini 3.8 Flash running inside Google Antigravity. Google's announcement notes that the choice of agent setup shaped outcomes on its own, separate from which underlying model was doing the work.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now