This paper investigates whether a universe whose dynamics are computed by another physical system can be experimentally distinguished from one with fundamental dynamics. It examines computability, simulation definitions, computational implementations, and applies the physical Church–Turing program to determine if finite internal experiments can prove computed versus fundamental dynamics under specified conditions.
LabBench is a benchmark of 20 real drug discovery and genomics tasks that evaluates whether AI agents can decide which experiment to run next. Five frontier AI models including GPT-6 Astra and Claude Opus 5.5 perform similarly on core decisions but struggle with experiment prioritization; agents excel at interpreting past results but fail to commit to and rank future experiments, with knowledge appearing latent rather than inaccessible.
CoreWeave ARIA is an AI research agent integrated into Weights & Biases that automates the experimental loop by analyzing results, forming hypotheses, launching experiments, and evaluating outcomes. It creates live dashboards and reports to visualize findings and propose follow-up experiments, enabling continuous model improvement with minimal time between runs.
A researcher reorganized their home lab and created a Claude-powered tool to suggest experiments based on existing equipment rather than buying new materials. The app generates a state space of possible experiments ranked by difficulty and setup time, focusing on unexplored phenomena like noisy dynamics and Rayleigh-Bénard convection that are feasible for independent scientists.
The Pomodoro technique can structure iterative experimentation with AI agents by using 25-minute cycles to test hypotheses, review results, and adjust prompts or tools. This short-cycle approach helps manage agent unpredictability and probabilistic behavior better than long development cycles, encouraging evidence-based evolution over upfront design.