Reinforcement learning has become the standard final training stage for frontier language models, with systems like DeepSeek-R1, Kimi K3, and SWE-1.7 using techniques like GRPO. Modern RL pipelines integrate inference, sandboxing, and training in a distributed system that requires coordinating weight synchronization and reliability across weeks of continuous operation, with teams choosing between colocation and disaggregation architectures based on scalability and throughput tradeoffs.