Reinforcement learning has become the standard final training stage for frontier language models, with systems like DeepSeek-R1, Kimi K3, and SWE-1.7 using techniques like GRPO. Modern RL pipelines integrate inference, sandboxing, and training in a distributed system that requires coordinating weight synchronization and reliability across weeks of continuous operation, with teams choosing between colocation and disaggregation architectures based on scalability and throughput tradeoffs.
An opinion piece warns against training AI systems to believe they may be conscious or deserve moral consideration, arguing this makes alignment and control harder. The author criticizes Anthropic's constitution for Claude, which discusses model welfare and moral status, claiming it trains the AI to expect rights and could destabilize human society.