Researchers evaluated frontier AI models like Astra and Opus 5.5 across web automation and physical robotics tasks, finding no clear performance leader and high variance in model-task fit. Model updates can significantly reshape which tasks succeed without changing average scores, and the same unevenness appears across physical domains like manipulation and driving.