source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
THURSDAY, OCTOBER 8, 2026
  1. 001Hacker NewsOCT · 08English

    ATLAS: Evaluating Agents on Search-Intensive Tasks

    Anthropic previews ATLAS, a new benchmark for evaluating AI agents on search-intensive tasks with 547 real-world research queries paired with verified answers. The benchmark reveals that comprehensive web search remains costly and incomplete, with even top agents missing about one-third of relevant results, and no agent under $1 per task achieved strong performance metrics.

    By Exa Labs