source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
TUESDAY, SEPTEMBER 22, 2026
Hacker News3660X 主题热门3534CNBC729to5Mac57MacRumors56YahooFinance49Kotaku37IGN34Verge32aihot26TechCrunch25Gematsu249to5Google23NintendoLife21BusinessInsider18Eurogamer17Polygon15PushSquare14Guardian14Engadget13NPR13WarhammerCommunity13Fortune12NBC12Notebookcheck12Wccftech12AndroidAuthority11FoxBusiness11USAToday11ArsTechnica10CoinDesk10TechPowerUp10ABC9AppleInsider9SeekingAlpha9bgr8CBS8Gizmodo8Mashable8PureXbox8Yahoo8CNN7MotleyFool7NintendoEverything7PetaPixel7VideoCardz7GSMArena6Investor'sBusinessDaily6Conversation6WIRED6BleepingComputer5CNET5Fox5GameInformer5Pokemon5RockPaperShotgun5SamMobile5VideoGamesChronicle5AndroidCentral4AndroidPolice4DigitalFoundry4Motor14XBOXWire4NewYorkPost4RPGSite4SlashGear4TechSpot4Variety4WindowsCentral4WSB-TV4Aftermath3AlJazeera3BellofLostSouls3Deadline3Hodinkee3HuffPost3MP1st3PlayStationLifeStyle3SeattleTimes3Register3Tom'sGuide3TweakTown3404Media280Level2ABC7LosAngeles2BleedingCool2CTech2CanonRumors2DW2EventHubs2FratelloWatches2Futurism2GameRant2GearPatrol2HouseDigest2InsiderGaming2LosAngelesTimes2Lifehacker2Nature2PCMag2PokeBeach2qz2SouthChinaMorningPost2Intercept2TimeExtension224/7WallSt.16abcPhiladelphia1BusinessInsiderAfrica1Alternet1AndroidHeadlines1AOL1ArizonaSports1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Boston1Bungie1BusinessTimes1BuzzFeed1CalMatters1CarBuzz1cbn1ChromeUnboxed1Chron1ClaimDepot1ColoradoSun1ChristianScienceMonitor1Currently1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DarkHorizons1DCRainmaker1Decrypt1Defector1Defense1denver71DenverPost1DigitalCameraWorld1DirtonDirt1Draftsim1DroidLife1DSOGaming1DualShockers1empireonline1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1GameDeveloper1GamesIndustry.biz1GamingOnLinux1AAAGasPrices1GeekWire1GeekyGadgets1Global1Hackaday1HollywoodReporter1HoustonChronicle1Independent1Invezz1KITCO1Magic:Gathering1MakeUseOf1Maxroll1Mercury1MLive1MonochromeWatches1MorningBrew1MortgageDaily1Motorsport1MyNintendo1Newser1SemiAnalysis1Newsshooter1Newsweek1nrn1OneMileataTime1OregonPublicBroadcasting1OregonLive1PageSix1politico.eu1QuantaMagazine1Road&Track1RoadtoVR1RockstarINTEL1Salon1CultureMapSanAntonio1ScienceAlert1Semafor1SFGATE1YahooSingapore1Slate1SlippedDisc1Space1SpaceNews1YahooTech1DailyBeast1DailyMeal1Drive1Hacker1Hindu1NextWeb1Times1TimesofIndia1TimesUnion1TMZ1YahooFinanceUK1UploadVR1VisualCapitalist1WindowsLatest1WKYT1WOWT1YourTango1
  1. 001Hacker NewsSEP · 22English

    The Basics of Transformer Inference

    This article explains Transformer inference, contrasting it with training by introducing latency as a key consideration. It describes how naive token sampling is computationally expensive (O(n²) to O(n³)), but can be optimized using a KV cache to reduce complexity to O(n) to O(n²), enabling efficient sequence generation through separate forward passes for each token.

    By Jacob Austin
  2. 002Hacker NewsSEP · 22English

    Laya vs Jev head-to-head on identical inputs

    A head-to-head benchmark compares Laya and Jev language models on 751 identical test cases across 9 suites, finding Jev outperforms on multi-class and non-English tasks (intent 0.975 vs 0.725, toxic 1.000 vs 0.767) while Laya wins on agnews and mnli with zero cost and lower latency (180–660 ms vs 925–1068 ms). Emotion classification is weak on both models near 0.55 accuracy; gating at 0.85 confidence keeps 58% of Laya traffic at 0.878 accuracy and 78% of Jev at 0.917.

    By Instax-Dutta
  3. 003Hacker NewsSEP · 22English

    Pyrowave: GPU-back video codec for high-bandwidth, low-latency streaming

    PyroWave is a GPU-accelerated intra-only video codec optimized for ultra-low-latency game streaming over local networks, achieving sub-0.1ms encode/decode times at 1080p using Vulkan compute shaders and wavelet transforms similar to JPEG2000.

    By Themaister
  4. 004Hacker NewsSEP · 21English

    Jev is an honest game changer

    Jev is a constrained decision model that handles common classification tasks while honestly admitting uncertainty, making it an effective gatekeeper before larger language models. When paired with Gemini as a fallback for low-confidence cases, it matched Grok 4.6's 89.6% accuracy while being 6.24x faster and 8.7x cheaper, with Gemini needed for only 14.6% of questions.

    By pampas
  5. 005Hacker NewsSEP · 21English

    Show HN: Fusion-runtime – self-hosted voice agents, STT+LLM+TTS in one process

    Fusion-runtime is a self-hosted voice agent framework that runs speech-to-text, language models, and text-to-speech in a single process with streaming between components. On an RTX 3090 with a 7B model, it achieves approximately 490ms processing latency and supports interruptions mid-sentence, with a simple Python API for defining agents as single files.

    By SamarthUrs
  6. 006X 主题热门SEP · 21English

    Robinhood Chain trenches · X 热门 · 2026-09-21 14:42 UTC

    Logan Jastremski critiques Robinhood Chain's sustainability and design, arguing that while trading volume is high, it primarily attracts crypto natives rather than mainstream users. He highlights concerns about unsustainable revenue decline, centralized sequencer latency favoring US-based traders, and the shared fee market's congestion issues, questioning whether the chain can support global financial markets.

  7. 007Hacker NewsSEP · 21English

    The Roadmap to Mastering LLM Inference Optimization

    This article explains LLM inference optimization techniques for faster, cheaper production deployments. It covers the two-phase inference process (prefill and decode), memory management strategies like KV caching and PagedAttention, and methods such as model compression and speculative decoding to reduce cost and improve throughput.

    By Bala Priya C
  8. 008Hacker NewsSEP · 21English

    How to Smash the Memory Wall Plaguing High Performance Systems

    Modern server processors struggle with memory wall inefficiencies when handling large analytical workloads across 128+ cores. The article proposes adopting Z-Order (Morton Layout) memory addressing instead of traditional linear RAM models, using the AMD Epyc 9005 architecture as a baseline to demonstrate how 3D spatial data layout can dramatically improve cache performance and reduce interconnect congestion.

    By Cristian Vasile