Anthropic previews ATLAS, a new benchmark for evaluating AI agents on search-intensive tasks with 547 real-world research queries paired with verified answers. The benchmark reveals that comprehensive web search remains costly and incomplete, with even top agents missing about one-third of relevant results, and no agent under $1 per task achieved strong performance metrics.