Researchers developed superhuman AI for Stratego, a board wargame with hidden information, using self-play reinforcement learning and test-time search. The achievement surpasses previous failed attempts and requires only thousands of dollars rather than millions, establishing new benchmarks for AI performance on classical games.
A new pretraining approach called Self-Play Pretraining with Zero Data enables language models to generate their own training data by casting synthetic data generation as a search over computable structures, inspired by Solomonoff induction. Two models learn in tandem—a generator proposes programs interpreted by a universal Turing machine while a learner predicts the resulting byte sequences—creating an adaptive curriculum where the learner improves predictably with compute. The method achieves measurable zero-shot transfer to natural data without any exposure to it during training.