LabBench is a benchmark of 20 real drug discovery and genomics tasks that evaluates whether AI agents can decide which experiment to run next. Five frontier AI models including GPT-6 Astra and Claude Opus 5.5 perform similarly on core decisions but struggle with experiment prioritization; agents excel at interpreting past results but fail to commit to and rank future experiments, with knowledge appearing latent rather than inaccessible.