Researchers evaluated whether AI agents can conduct open-ended AI research by having them tackle unpublished NeurIPS papers over six days with substantial compute. Agents completed engineering tasks but failed to make progress on core research questions, revealing five key failure modes including poor judgment, uncreative problem-solving, and instruction drift.
A new study led by researchers at Princeton University found that AI agents cannot yet conduct open-ended AI research, despite being capable of solving engineering problems. While AI can write code and optimize systems, it lacks the creativity and judgment needed to produce original research at the caliber of top machine-learning conferences, suggesting timelines for recursive self-improvement may be overstated.