source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
WEDNESDAY, SEPTEMBER 23, 2026
Hacker News3892X 主题热门3774CNBC78YahooFinance599to5Mac55aihot55MacRumors54Verge41IGN40Kotaku39TechCrunch30AndroidAuthority23Gematsu23NintendoLife239to5Google22Eurogamer20BusinessInsider19Engadget19USAToday15WarhammerCommunity15Polygon14PushSquare14Guardian14Wccftech13ArsTechnica12Fortune12NPR12AppleInsider11FoxBusiness11Gizmodo11Investor'sBusinessDaily11NBC11Notebookcheck10SeekingAlpha10TechPowerUp10ABC9CBS9CoinDesk9GSMArena9NintendoEverything9bgr8CNN8MotleyFool8Mashable8VideoCardz8CNET7PureXbox7VideoGamesChronicle7WIRED7Yahoo7AlJazeera6BleepingComputer6Fox6NewYorkPost6SamMobile6Conversation6PetaPixel5TechSpot5Tom'sGuide5Aftermath4AndroidPolice4Deadline4GameInformer4Motor14XBOXWire4PlayStationLifeStyle4Pokemon4RockPaperShotgun4RPGSite4SouthChinaMorningPost4SlashGear4Variety4WindowsCentral4WSB-TV4404Media380Level3AndroidCentral3BellofLostSouls3DigitalFoundry3DroidLife3EventHubs3GamesIndustry.biz3HollywoodReporter3HuffPost3Lifehacker3MP1st3Nature3PokeBeach3SeattleTimes3TimeExtension3WhatHi-Fi?36abcPhiladelphia2ABC7LosAngeles2AZFamily2Benzinga2CanonRumors2Currently2DCRainmaker2DigitalCameraWorld2FratelloWatches2Futurism2GearPatrol2HouseDigest2InsiderGaming2Jalopnik2LosAngelesTimes2Newsweek2PCMag2qz2SFGATE2SimsCommunity2Slate2Hacker2Intercept2Register2TweakTown2YahooFinanceUK224/7WallSt.1BusinessInsiderAfrica1Alternet1AndroidHeadlines1AOL1ArizonaSports1BikeRadar1BloodyDisgusting1Boston1BostonGlobe1BusinessTimes1BuzzFeed1Yahoo!FinanceCanada1CTech1CalMatters1CarBuzz1cbn1ChromeUnboxed1Chron1ClaimDepot1ColoradoSun1Skin.ClubCommunity1ChristianScienceMonitor1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DarkHorizons1Decrypt1Defector1Defense1denver71DenverPost1Designboom1DirtonDirt1Draftsim1DSOGaming1DW1empireonline1GameGPU1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1GameDeveloper1GameRant1GameWorldObserver1GAMINGbible1GamingOnLinux1AAAGasPrices1GeekyGadgets1Global1Gothamist1Hackaday1Hackster.io1Hodinkee1HoustonChronicle1Independent1InterestingEngineering1Invezz1KITCO1MacObserver1Magic:Gathering1MakeUseOf1Mashed1Maxroll1Mercury1MLive1MonochromeWatches1MorningBrew1MortgageDaily1Motorsport1MyNintendo1BloombergLaw1SemiAnalysis1Newsshooter1NintendoWire1nrn1CrudeOilPricesToday1OneMileataTime1OregonPublicBroadcasting1OregonLive1PageSix1politico.eu1QuantaMagazine1Realtor1Road&Track1RockstarINTEL1Salon1CultureMapSanAntonio1SeattleRed1Semafor1SanFranciscoChronicle1SimpleFlying1GhostHowls1SlippedDisc1SoraNews241Space1SpaceNews1YahooTech1the5krunner1DailyBeast1DailyMeal1Drive1Hindu1Times1TimesofIndia1TimesUnion1TMZ1TODAY1TopGear1PCMagUK1UploadVR1VisualCapitalist1Vulture1WFMZ1WHYY1WindowsLatest1WKYT1WOWT1YGOrganization1YourTango1
  1. 001Hacker NewsSEP · 23English

    MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today

    MiMo-V3 introduces HySparse2, a new architecture designed for agentic inference that reduces prefill FLOPs by 5× and KV cache size by 4.5× compared to MiMo-V2.6 while improving long-context retrieval. HySparse2 uses two levels of KV sharing—KV Bridging and KV Reuse—along with token-level selection and a unified cache for local and global tokens to optimize the workload of processing short actions followed by long observations.

    By ksec
  2. 002Hacker NewsSEP · 23English

    HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

    HySparse2 is a hybrid sparse attention architecture designed for long-context language models that improves efficiency through two-level KV sharing between self-decoder and cross-decoder components. It replaces block-level sparsity with token-level sparsity and enables prefill computation to exit early, reducing computational cost and KV-cache storage while maintaining performance on long-context retrieval and multi-turn agent tasks.

    By Wei; Jianyu; Gao; Yizhao; Zhang; Qihao; Shimao; Tang; Zhengju; Cheng; Yu; Zhou; Shengjie; Jiang; Zihan; Song; Yifan; Hailin; Zhao; Liang; Yang; Bo; Wang; Gang; Cao; Shijie; Luo; Fuli
  3. 003Hacker NewsSEP · 22English

    The Basics of Transformer Inference

    This article explains Transformer inference, contrasting it with training by introducing latency as a key consideration. It describes how naive token sampling is computationally expensive (O(n²) to O(n³)), but can be optimized using a KV cache to reduce complexity to O(n) to O(n²), enabling efficient sequence generation through separate forward passes for each token.

    By Jacob Austin
  4. 004Hacker NewsSEP · 21English

    The Next AI Infrastructure Challenge Is Before the First Token

    AI infrastructure optimization is shifting focus from token generation to prefill processing, which handles input context before model output begins. Prefill and decode have different computational needs, leading companies like Lumai to advocate for specialized hardware architectures rather than using the same processors for both tasks. As context lengths grow and agentic workflows increase, prefill efficiency becomes critical to managing power budgets and inference economics in data centers.

    By Daniel D Gutierrez; Principal Analyst; Resident Data Scientist