Google released Android Bench 2.0, a benchmark for evaluating AI models on complex Android development tasks requiring multiple days to complete, such as building apps from scratch and porting cross-platform applications. The new version uses continuous scoring based on functionality, visual fidelity, and regression avoidance, with GPT-6 Astra achieving the highest pass rate at 28%. Testing revealed that models excel at writing new code and deterministic transformations but struggle with refactoring, runtime validation, and unfamiliar libraries.