Cohere introduces a serving engine for North Mini Code using a decode megakernel, achieving 1.25–1.41× speedup over vLLM on H100 GPUs by consolidating multiple kernel launches into a single persistent kernel that runs the entire forward pass, eliminating GPU idle time during autoregressive decoding.