A study audits the SWE-bench leaderboard for coding agents, finding that top entries have converged with highly overlapping success sets, making small score differences unreliable for ranking. Statistical tests show most adjacent top-thirty pairs cannot be significantly distinguished, suggesting leaderboard positions don't establish clear ordering. The authors propose reporting model-scaffold-specific results and resolution metrics instead of interpreting marginal aggregate gaps as meaningful rank differences.