Inferact released TPU megakernels for Google's TPU v7 that achieve over 700 tokens/second with the Kimi K3 model, significantly outperforming NVIDIA's GB200 GPU. The megakernels leverage TPU's large on-chip memory and explicit asynchronous programming to keep memory bandwidth utilization high during inference, delivering 1.4 to 2× better decode throughput at small batch sizes.