source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
TUESDAY, SEPTEMBER 22, 2026
Hacker News3699X 主题热门3577CNBC739to5Mac57MacRumors56YahooFinance50Kotaku37IGN34Verge34aihot27TechCrunch25Gematsu249to5Google23NintendoLife21BusinessInsider18Eurogamer17Polygon15NPR14PushSquare14Guardian14Engadget13WarhammerCommunity13Fortune12NBC12Notebookcheck12Wccftech12AndroidAuthority11FoxBusiness11USAToday11ArsTechnica10CoinDesk10TechPowerUp10ABC9AppleInsider9SeekingAlpha9bgr8CBS8Gizmodo8Mashable8PureXbox8Yahoo8CNN7MotleyFool7NintendoEverything7PetaPixel7Conversation7VideoCardz7GSMArena6Investor'sBusinessDaily6SamMobile6WIRED6BleepingComputer5CNET5Fox5GameInformer5Pokemon5RockPaperShotgun5VideoGamesChronicle5AndroidCentral4AndroidPolice4DigitalFoundry4Motor14XBOXWire4NewYorkPost4RPGSite4SlashGear4TechSpot4Variety4WindowsCentral4WSB-TV4Aftermath3AlJazeera3BellofLostSouls3Deadline3Hodinkee3HuffPost3MP1st3PlayStationLifeStyle3SeattleTimes3Register3Tom'sGuide3TweakTown3404Media26abcPhiladelphia280Level2ABC7LosAngeles2BleedingCool2CTech2CanonRumors2DW2EventHubs2FratelloWatches2Futurism2GameRant2GamesIndustry.biz2GearPatrol2HouseDigest2InsiderGaming2LosAngelesTimes2Lifehacker2Nature2PCMag2PokeBeach2qz2SouthChinaMorningPost2Intercept2TimeExtension224/7WallSt.1BusinessInsiderAfrica1Alternet1AndroidHeadlines1AOL1ArizonaSports1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Boston1Bungie1BusinessTimes1BuzzFeed1CalMatters1CarBuzz1cbn1ChromeUnboxed1Chron1ClaimDepot1ColoradoSun1ChristianScienceMonitor1Currently1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DarkHorizons1DCRainmaker1Decrypt1Defector1Defense1denver71DenverPost1DigitalCameraWorld1DirtonDirt1Draftsim1DroidLife1DSOGaming1DualShockers1empireonline1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1GameDeveloper1GamingOnLinux1AAAGasPrices1GeekWire1GeekyGadgets1Global1Hackaday1HollywoodReporter1HoustonChronicle1Independent1Invezz1Jalopnik1KITCO1Magic:Gathering1MakeUseOf1Mashed1Maxroll1Mercury1MLive1MonochromeWatches1MorningBrew1MortgageDaily1Motorsport1MyNintendo1Newser1SemiAnalysis1Newsshooter1Newsweek1nrn1OneMileataTime1OregonPublicBroadcasting1OregonLive1PageSix1politico.eu1QuantaMagazine1Road&Track1RoadtoVR1RockstarINTEL1Salon1CultureMapSanAntonio1ScienceAlert1Semafor1SFGATE1YahooSingapore1Slate1SlippedDisc1Space1SpaceNews1YahooTech1DailyBeast1DailyMeal1Drive1Hacker1Hindu1NextWeb1Times1TimesofIndia1TimesUnion1TMZ1YahooFinanceUK1UploadVR1VisualCapitalist1WindowsLatest1WKYT1WOWT1YourTango1
  1. 001Hacker NewsSEP · 22English

    The Basics of Transformer Inference

    This article explains Transformer inference, contrasting it with training by introducing latency as a key consideration. It describes how naive token sampling is computationally expensive (O(n²) to O(n³)), but can be optimized using a KV cache to reduce complexity to O(n) to O(n²), enabling efficient sequence generation through separate forward passes for each token.

    By Jacob Austin
  2. 002Hacker NewsSEP · 21English

    The Next AI Infrastructure Challenge Is Before the First Token

    AI infrastructure optimization is shifting focus from token generation to prefill processing, which handles input context before model output begins. Prefill and decode have different computational needs, leading companies like Lumai to advocate for specialized hardware architectures rather than using the same processors for both tasks. As context lengths grow and agentic workflows increase, prefill efficiency becomes critical to managing power budgets and inference economics in data centers.

    By Daniel D Gutierrez; Principal Analyst; Resident Data Scientist
  3. 003Hacker NewsSEP · 16English

    Prefill a 284B model on Nvidia. Decode it on Apple Silicon. Over plain 10GbE

    A system successfully prefills a 284B parameter DeepSeek-V4-Flash model on NVIDIA DGX hardware and decodes it on Apple Silicon Mac Studio over standard 10GbE ethernet, achieving 1.5x to 3.7x speedup on prompts up to 241K tokens by computing the decoder's cache on the prefill machine rather than transferring incompatible KV cache formats.

    By Chadhurley
  4. 004Hacker NewsSEP · 16English

    Deep Seek v4.1 M5 Max at 17 tokens/s

    Deep Seek v4.1 M5 Max achieves 17 tokens/s by optimizing mixture-of-experts model execution from SSDs on a 128GB laptop. The system reads only routed experts (187 of 384 per layer) instead of all experts, achieving 1.73× speedup on prefill; adding multiple drives reduces read latency rather than increasing bandwidth, with time-to-first-token improving from 31.5s (baseline) to 18.3s (one drive) to 11.7s (three drives).

    By Argonautlabsai