Anthropic introduced sudo L7, a benchmark measuring staff-level software engineering judgment beyond basic coding ability. It evaluates 60 tasks from real production codebases across security, reliability, and architecture, grading agents on functional correctness, architectural decisions, risk awareness, and communication—not just whether tests pass. Claude Opus 5.5 and Sonnet 5.5 completed roughly 45% of tasks, revealing larger differences in engineering behavior beyond overall scores.