A benchmark comparing NVFP4 and MXFP4 quantization formats on NVIDIA B200 GPUs shows NVFP4 delivers up to 8% faster decode performance at small batch sizes when running Qwen3-32B through vLLM, with differences disappearing at larger batches due to kernel implementation variations rather than memory bandwidth constraints.