Researchers evaluated frontier AI models like Astra and Opus 5.5 across web automation and physical robotics tasks, finding no clear performance leader and high variance in model-task fit. Model updates can significantly reshape which tasks succeed without changing average scores, and the same unevenness appears across physical domains like manipulation and driving.
Researchers evaluated frontier AI models like Astra and Opus 5.5 across web automation and physical robotics tasks, finding no clear performance leader and high variance in model-task fit. Model updates can reshape performance profiles without changing average scores, and individual task success rates vary significantly even within single benchmarks.