source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
SATURDAY, OCTOBER 10, 2026
  1. 001Hacker NewsOCT · 10English

    NVFP4 vs. MXFP4 Decode Benchmark

    A benchmark comparing NVFP4 and MXFP4 quantization formats on NVIDIA B200 GPUs shows NVFP4 delivers up to 8% faster decode performance at small batch sizes when running Qwen3-32B through vLLM, with differences disappearing at larger batches due to kernel implementation variations rather than memory bandwidth constraints.

    By Cezar Cocu
  2. 002Hacker NewsOCT · 09English

    Local LLM Inference at Scale with vLLM

    vLLM is a full-featured serving engine for self-hosting open-weight language models at scale, capable of handling thousands of requests through innovations like PagedAttention and continuous batching. The post evaluates vLLM's performance characteristics on NVIDIA hardware, demonstrating how bandwidth constraints, quantization strategies, and model architecture choices affect throughput for local LLM inference.

    By Bruno Gonçalves
  3. 003Hacker NewsOCT · 09English

    Show HN: Long term Memory and 50M token window for LLM

    Researchers tested galahad-kv, a memory layer that caches key-value states of LLM blocks on encrypted NVMe disk to enable 50-million-token context windows. Loading cached blocks proved 2.8–4.3x faster and 8.8–12.3x more energy-efficient than recomputing them, with both Gemma 12B and 31B models accurately recalling facts from millions of tokens earlier.

    By Sietse Schelpe
  4. 004Hacker NewsOCT · 09English

    Long-Term Memory for AI:50M-Token Window,Is Faster,Cheaper Than Recompute

    Researchers demonstrate a memory layer that extends AI language models to handle 50-million-token contexts by storing and retrieving key-value states from encrypted disk storage, achieving 2.8-4.3x faster loading and 8.8-12.3x lower GPU energy use compared to recomputation. Testing on Gemma models shows accurate retrieval of facts from millions of tokens earlier with no hallucinations.

    By Schelpe; Sietse
  5. 005Hacker NewsOCT · 08English

    Vosti: Specifying, Implementing, and Verifying Deterministic LLM Inference

    Vosti is a formally verified LLM inference engine that ensures deterministic outputs by producing bitwise-identical logits across different execution variations. The system addresses limitations in production systems like vLLM and SGLang by formalizing deterministic inference specifications and proving correctness through decomposed proofs at the engine and GPU kernel boundaries.

    By Qin; Jianxing; Du; Alexander; Zhang; Danfeng; Lentz; Matthew; Zhuo; Danyang