LoopCD is a training-free contrastive decoding framework for looped Transformers that improves token prediction by contrasting final outputs with earlier recurrent passes. The method achieves substantial performance gains—raising AIME pass rates from 61.88% to 73.33% and HumanEval from 22.56% to 31.71%—while reducing inference compute by 22.5% to 48.2% through fewer required loops.
This paper introduces Loop Scaling Laws, the first scaling framework that jointly models recurrence and sparsity in looped Mixture-of-Experts transformers. The work demonstrates that combining recurrence and MoE sparsity delivers complementary efficiency gains: ~3x active-parameter efficiency from sparsity and ~2x total-parameter efficiency from recurrence on reasoning tasks, with practical validation at trillion-token scale.