source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
TUESDAY, SEPTEMBER 22, 2026
Hacker News3796X 主题热门3673CNBC779to5Mac57MacRumors57YahooFinance52Kotaku40Verge36IGN35aihot30TechCrunch259to5Google24Gematsu24NintendoLife23BusinessInsider19Eurogamer18Polygon15PushSquare15WarhammerCommunity15Engadget14NPR14Guardian14AndroidAuthority13NBC13Fortune12Notebookcheck12USAToday12Wccftech12FoxBusiness11AppleInsider10ArsTechnica10CoinDesk10TechPowerUp10ABC9Mashable9SeekingAlpha9bgr8CBS8CNN8Gizmodo8PureXbox8Yahoo8MotleyFool7Investor'sBusinessDaily7NintendoEverything7PetaPixel7Conversation7VideoCardz7GSMArena6SamMobile6VideoGamesChronicle6WIRED6BleepingComputer5CNET5Fox5GameInformer5Pokemon5RockPaperShotgun5AndroidCentral4AndroidPolice4DigitalFoundry4Motor14XBOXWire4NewYorkPost4RPGSite4SlashGear4TechSpot4Variety4WindowsCentral4WSB-TV4Aftermath3AlJazeera3BellofLostSouls3Deadline3Hodinkee3HuffPost3MP1st3PlayStationLifeStyle3SouthChinaMorningPost3SeattleTimes3Register3Tom'sGuide3TweakTown3404Media26abcPhiladelphia280Level2ABC7LosAngeles2BleedingCool2CTech2CanonRumors2DCRainmaker2DigitalCameraWorld2DroidLife2DW2EventHubs2FratelloWatches2Futurism2GameRant2GamesIndustry.biz2GearPatrol2HouseDigest2InsiderGaming2LosAngelesTimes2Lifehacker2Nature2PCMag2PokeBeach2qz2Intercept2TimeExtension224/7WallSt.1BusinessInsiderAfrica1Alternet1AndroidHeadlines1AOL1ArizonaSports1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Boston1Bungie1BusinessTimes1BuzzFeed1CalMatters1CarBuzz1cbn1ChromeUnboxed1Chron1ClaimDepot1ColoradoSun1ChristianScienceMonitor1Currently1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DarkHorizons1Decrypt1Defector1Defense1denver71DenverPost1DirtonDirt1Draftsim1DSOGaming1DualShockers1empireonline1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1GameDeveloper1GamingOnLinux1AAAGasPrices1GeekWire1GeekyGadgets1Global1Hackaday1HollywoodReporter1HoustonChronicle1Independent1Invezz1Jalopnik1KITCO1Magic:Gathering1MakeUseOf1Mashed1Maxroll1Mercury1MLive1MonochromeWatches1MorningBrew1MortgageDaily1Motorsport1MyNintendo1Newser1SemiAnalysis1Newsshooter1Newsweek1nrn1CrudeOilPricesToday1OneMileataTime1OregonPublicBroadcasting1OregonLive1PageSix1politico.eu1QuantaMagazine1Road&Track1RoadtoVR1RockstarINTEL1Salon1CultureMapSanAntonio1ScienceAlert1Semafor1SFGATE1YahooSingapore1Slate1SlippedDisc1Space1SpaceNews1YahooTech1the5krunner1DailyBeast1DailyMeal1Drive1Hacker1Hindu1NextWeb1Times1TimesofIndia1TimesUnion1TMZ1YahooFinanceUK1UploadVR1VisualCapitalist1WindowsLatest1WKYT1WOWT1YourTango1
  1. 001Hacker NewsSEP · 22English

    GPU is starving – LLM host dispatch at 191k steps/s on 1 vCPU

    Floria is a high-throughput serving engine for large language models that eliminates GPU starvation by replacing Python-based scheduling with native hardware dispatching. Running on a single vCPU, it achieves 191k tokens/sec and 100% GPU utilization, compared to conventional systems like vLLM that leave GPUs idle 25-40% of the time due to host scheduling latency.

    By Cortexlab
  2. 002Hacker NewsSEP · 21English

    Halo: Post-train LLMs 3x faster than TRL and Megatron

    Halo is a new post-training framework for open-source models that achieves 2.8x higher throughput than TRL while using less peak memory and maintaining HuggingFace compatibility.

    By ovyan
  3. 003Hacker NewsSEP · 21English

    The Roadmap to Mastering LLM Inference Optimization

    This article explains LLM inference optimization techniques for faster, cheaper production deployments. It covers the two-phase inference process (prefill and decode), memory management strategies like KV caching and PagedAttention, and methods such as model compression and speculative decoding to reduce cost and improve throughput.

    By Bala Priya C
  4. 004Hacker NewsSEP · 20English

    Using system-one models inside high-throughput data pipelines

    A system-one model called Jev significantly improves entity resolution in high-throughput data pipelines, reducing costs by 99.56% and increasing throughput 7.35× while maintaining near-baseline accuracy when integrated into multi-stage review workflows. The authors demonstrate this on Ohio campaign finance data, where Jev resolves entities across donors, committees, and companies by adjudicating bundles of equivalent evidence.

    By Southbridge AI; Hrishi Olickel
  5. 005Hacker NewsSEP · 18English

    Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

    ShapeLearn released optimized GGUF quantizations of Qwen 3.8 27B, with full models outperforming their earlier Lite versions. GPU-5 is recommended for best quality-speed tradeoff at 99.63% of BF16 performance, while GPU-4 offers competitive results in smaller 11.0 GB size. Speculative decoding with MTP or DFlash2 further improves throughput across all tested GPUs.

    By syntaxing
  6. 006Hacker NewsSEP · 17English

    Union Alpha is Unbiased's Pareto

    Union Alpha is Unbiased's Pareto, a multimodal composite model hosted by OpenRouter with 262,144 token context window, priced at $2.50/M input and $7.50/M output tokens. The model delivers 34 tokens/second throughput and 4.52 seconds latency, supporting function calling and JSON output for research, coding, and agentic workflows.

    By algoth1
  7. 007Hacker NewsSEP · 16English

    Characterizing ~375 GB Kimi K2.5 on a 128 GB Ryzen AI MAX+ 395 PC

    A systems characterization study examines executing the ~375 GB Kimi K2.5 model on a single 128 GB AMD Ryzen AI MAX+ 395 PC using storage-backed expert caching. The bounded locality cache achieved 7.7% hit rate, reducing expert traffic by ~70 GB and decreasing package energy from 117.50 J to 113.47 J per generated token while maintaining ~0.438 tokens/s throughput.

    By Hutchinson; Nigel
  8. 008Hacker NewsSEP · 16English

    How VoltDB Works

    VoltDB is a distributed database that optimizes performance by partitioning data and stored procedures across multiple sites, enabling parallel query execution without traditional locking overhead. Transactions are serialized within partitions for consistency, while multi-partition operations use coordination to maintain throughput and integrity.

    By ronfriedhaber
  9. 009Hacker NewsSEP · 16English

    Dense vs. Moe Models: Active Parameters, Throughput, and When to Choose Each

    Dense and Mixture-of-Experts (MoE) model architectures differ fundamentally in parameter activation: dense models activate all parameters for every token, while MoE models route each token through only a subset of expert networks. The choice between them depends on deployment constraints like throughput, memory cost, and serving complexity rather than raw parameter count alone.

    By Elizabeth Goodman
  10. 010Hacker NewsSEP · 15English

    Show HN: Self-adjusting vLLM at production scale

    Rivvr automates vLLM deployment at production scale by automatically tuning configurations, monitoring metrics, and adjusting cluster topology to maintain latency and throughput SLO targets while reducing AWS costs by 40-70%.

    By dmitriisn