A benchmark comparing NVFP4 and MXFP4 quantization formats on NVIDIA B200 GPUs shows NVFP4 delivers up to 8% faster decode performance at small batch sizes when running Qwen3-32B through vLLM, with differences disappearing at larger batches due to kernel implementation variations rather than memory bandwidth constraints.
vLLM is a full-featured serving engine for self-hosting open-weight language models at scale, capable of handling thousands of requests through innovations like PagedAttention and continuous batching. The post evaluates vLLM's performance characteristics on NVIDIA hardware, demonstrating how bandwidth constraints, quantization strategies, and model architecture choices affect throughput for local LLM inference.
Researchers tested galahad-kv, a memory layer that caches key-value states of LLM blocks on encrypted NVMe disk to enable 50-million-token context windows. Loading cached blocks proved 2.8–4.3x faster and 8.8–12.3x more energy-efficient than recomputing them, with both Gemma 12B and 31B models accurately recalling facts from millions of tokens earlier.
Researchers demonstrate a memory layer that extends AI language models to handle 50-million-token contexts by storing and retrieving key-value states from encrypted disk storage, achieving 2.8-4.3x faster loading and 8.8-12.3x lower GPU energy use compared to recomputation. Testing on Gemma models shows accurate retrieval of facts from millions of tokens earlier with no hallucinations.
Vosti is a formally verified LLM inference engine that ensures deterministic outputs by producing bitwise-identical logits across different execution variations. The system addresses limitations in production systems like vLLM and SGLang by formalizing deterministic inference specifications and proving correctness through decomposed proofs at the engine and GPU kernel boundaries.