GoBench is a benchmark that measures how well frontier language models play 9×9 Go against KataGo opponents calibrated on an Elo ladder. The arXiv paper is scheduled for release on September 17, 2026.
Artificial Analysis's Intelligence Index, a widely-used AI model leaderboard, compresses diverse evaluation choices into a single intelligence score that obscures methodological trade-offs. The index weights agentic workloads heavily (34%), relies on saturated benchmarks like GPQA that cannot distinguish frontier models, lacks genuine coding benchmarks despite a 24% coding category, and features half its components using similar agent-execution patterns that may over-represent certain capabilities.