source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
SATURDAY, OCTOBER 10, 2026
Hacker News1395X 主题热门1249ChainCatcher101CNBC72PANews55Coinness53Channel News Asia46SouthChinaMorningPost29Crypto.news26YahooFinance26CoinDesk21Verge21Cointelegraph18Kotaku18aihot169to5Mac15IGN15Variety15MacRumors13Decrypt119to5Google10TechCrunch10AMBCrypto9DW9Eurogamer9TechPowerUp9NintendoLife8Guardian8BBC World7FoxBusiness7BitcoinMagazine6BusinessInsider6Engadget6NBC6Polygon6VideoCardz6Wccftech6ArsTechnica5Investor'sBusinessDaily5WarhammerCommunity5bgr4CBS4CNET4CNN4Gematsu4Gizmodo4Mashable4XBOXWire4Register4USAToday4Futurism3GSMArena3NYT3PokémonGOHub3PushSquare3VideoGamesChronicle3AlJazeera2BleepingComputer2DigitalFoundry2DroidLife2Euronews2MotleyFool2Fox2KSL2MyNintendo2Nature2Newser2Notebookcheck2NPR2SeattleTimes2SeekingAlpha2Hacker2WindowsCentral2WIRED26abcPhiladelphia1ABC7NewYork1ABC1AlineaInsightnewsletter1AndroidCentral1AndroidPolice1AppleInsider1NikkeiAsia1Bank of England1Barron's1Beebom1BloodyDisgusting1BostonGlobe1CFTC1ChromeUnboxed1Cleveland1CreativeBloq1DSOGaming1Federal Reserve1Fortune1FTC1GameDeveloper1GameFile1GameRant1GeekWire1HotHardware1Independent1InsiderGaming1InvenGlobal1KUTV1Lifehacker1MP1st1NBC5Chicago1NBCSports1MicrosoftSource1NintendoWire1CrudeOilPricesToday1OMG!Ubuntu1PaulKrugman1PennLive1Phoronix1Pokemon1PureXbox1RetractionWatch1Road&Track1RockPaperShotgun1RPGSite1ScienceAlert1SEC1SFGATE1YahooFinanceSingapore1Yahoo1SpaceNews1Syracuse1YahooTech1Hill1Outerhaven1Times1Tom'sGuide1TopGear1TweakTown1YahooUK1OutsideMagazine1VGChartz1EdZitron'sWhere'sYourEdAt1WolfStreet1Yahoo1YankoDesign1ZDNET1
  1. 001Hacker NewsOCT · 09English

    Local LLM Inference at Scale with vLLM

    vLLM is a full-featured serving engine for self-hosting open-weight language models at scale, capable of handling thousands of requests through innovations like PagedAttention and continuous batching. The post evaluates vLLM's performance characteristics on NVIDIA hardware, demonstrating how bandwidth constraints, quantization strategies, and model architecture choices affect throughput for local LLM inference.

    By Bruno Gonçalves