This paper investigates language drift—unusual, non-standard language in LLM reasoning chains—that emerges during reinforcement learning with verifiable reward (RLVR) post-training. The authors prove theoretically that RLVR permits unbounded language drift while supervised fine-tuning does not, and show empirically that drift occurs on novel reasoning tasks. They demonstrate that constraining language drift necessarily harms performance, presenting a fundamental trade-off in frontier LLM post-training.