ByteDance and Tsinghua researchers released DAPO, an open-source reinforcement learning system for large language models that achieves 50% accuracy on AIME 2024 using Qwen2.5-32B, outperforming previous state-of-the-art methods with fewer training steps. The system includes the Decoupled Clip and Dynamic sAmpling Policy Optimization algorithm, code infrastructure, and datasets for scalable LLM training.