source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
SUNDAY, OCTOBER 11, 2026
Hacker News1496X 主题热门1334ChainCatcher105CNBC72PANews56Coinness53Channel News Asia46SouthChinaMorningPost31YahooFinance27Crypto.news26CoinDesk21Verge21Kotaku19Cointelegraph18aihot17IGN169to5Mac15Variety15MacRumors149to5Google11Decrypt11TechCrunch11AMBCrypto10DW9Eurogamer9TechPowerUp9NintendoLife8Guardian8BBC World7FoxBusiness7BitcoinMagazine6BusinessInsider6Engadget6NBC6Polygon6VideoCardz6Wccftech6ArsTechnica5CNET5Investor'sBusinessDaily5XBOXWire5Register5WarhammerCommunity5bgr4CBS4CNN4Gematsu4Gizmodo4GSMArena4Mashable4NYT4USAToday4Fortune3Fox3Futurism3Notebookcheck3PokémonGOHub3PushSquare3VideoGamesChronicle3AlJazeera2BleepingComputer2DigitalFoundry2DroidLife2Euronews2MotleyFool2KSL2Lifehacker2MyNintendo2Nature2Newser2NPR2SeattleTimes2SeekingAlpha2Hacker2WindowsCentral2WIRED224/7WallSt.16abcPhiladelphia1ABC7NewYork1ABC1AlineaInsightnewsletter1AndroidCentral1AndroidPolice1AppleInsider1NikkeiAsia1Bank of England1Barron's1Beebom1BloodyDisgusting1BostonGlobe1CFTC1ChromeUnboxed1Cleveland1CreativeBloq1DSOGaming1Federal Reserve1FTC1GameDeveloper1GameFile1GameRant1GeekWire1HotHardware1Independent1InsiderGaming1InvenGlobal1KTLO1KUTV1MP1st1NBC5Chicago1NBCSports1MicrosoftSource1NintendoWire1CrudeOilPricesToday1OMG!Ubuntu1PaulKrugman1PennLive1Phoronix1Pocket-lint1Pokemon1PureXbox1RetractionWatch1Road&Track1RockPaperShotgun1RPGSite1ScienceAlert1SEC1SFGATE1YahooFinanceSingapore1Yahoo1SpaceNews1Syracuse1YahooTech1Hill1Outerhaven1Times1Tom'sGuide1TopGear1TweakTown1YahooUK1OutsideMagazine1VGChartz1EdZitron'sWhere'sYourEdAt1WolfStreet1Yahoo1YankoDesign1ZDNET1
  1. 001Hacker NewsOCT · 10English

    Show HN: Try Free Long-Term Memory for AI 50M-Token Window

    Galahad is a memory layer for AI models that caches KV tokens to avoid reprocessing, supporting up to 50M-token context windows. It integrates with llama.cpp, vLLM, and SGLang, offering 14× speedup, byte-exact retrieval, and encrypted persistent storage, with a free tier for single GPU usage.

    By Corbenic
  2. 002Hacker NewsOCT · 10English

    NVFP4 vs. MXFP4 Decode Benchmark

    A benchmark comparing NVFP4 and MXFP4 quantization formats on NVIDIA B200 GPUs shows NVFP4 delivers up to 8% faster decode performance at small batch sizes when running Qwen3-32B through vLLM, with differences disappearing at larger batches due to kernel implementation variations rather than memory bandwidth constraints.

    By Cezar Cocu
  3. 003Hacker NewsOCT · 09English

    Local LLM Inference at Scale with vLLM

    vLLM is a full-featured serving engine for self-hosting open-weight language models at scale, capable of handling thousands of requests through innovations like PagedAttention and continuous batching. The post evaluates vLLM's performance characteristics on NVIDIA hardware, demonstrating how bandwidth constraints, quantization strategies, and model architecture choices affect throughput for local LLM inference.

    By Bruno Gonçalves
  4. 004Hacker NewsOCT · 09English

    Show HN: Long term Memory and 50M token window for LLM

    Researchers tested galahad-kv, a memory layer that caches key-value states of LLM blocks on encrypted NVMe disk to enable 50-million-token context windows. Loading cached blocks proved 2.8–4.3x faster and 8.8–12.3x more energy-efficient than recomputing them, with both Gemma 12B and 31B models accurately recalling facts from millions of tokens earlier.

    By Sietse Schelpe
  5. 005Hacker NewsOCT · 09English

    Long-Term Memory for AI:50M-Token Window,Is Faster,Cheaper Than Recompute

    Researchers demonstrate a memory layer that extends AI language models to handle 50-million-token contexts by storing and retrieving key-value states from encrypted disk storage, achieving 2.8-4.3x faster loading and 8.8-12.3x lower GPU energy use compared to recomputation. Testing on Gemma models shows accurate retrieval of facts from millions of tokens earlier with no hallucinations.

    By Schelpe; Sietse
  6. 006Hacker NewsOCT · 08English

    Vosti: Specifying, Implementing, and Verifying Deterministic LLM Inference

    Vosti is a formally verified LLM inference engine that ensures deterministic outputs by producing bitwise-identical logits across different execution variations. The system addresses limitations in production systems like vLLM and SGLang by formalizing deterministic inference specifications and proving correctness through decomposed proofs at the engine and GPU kernel boundaries.

    By Qin; Jianxing; Du; Alexander; Zhang; Danfeng; Lentz; Matthew; Zhuo; Danyang