source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
WEDNESDAY, SEPTEMBER 16, 2026
Hacker News3634X 主题热门3523MacRumors78CNBC71YahooFinance649to5Mac59Kotaku44Verge43IGN339to5Google32aihot31Gematsu31NintendoLife30TechCrunch25Engadget24Eurogamer24BusinessInsider23Guardian20CNET15NBC15NPR15FoxBusiness14Fortune13Polygon13SeekingAlpha13bgr12Gizmodo12CBS11Wccftech11Investor'sBusinessDaily10Mashable10TechPowerUp10USAToday10WIRED10PushSquare9CNN8NintendoEverything8Notebookcheck8NewYorkPost8CrudeOilPricesToday8VideoGamesChronicle8ABC7ArsTechnica7Fox7GameInformer7WindowsCentral7BleepingComputer6AppleInsider5Deadline5GamesIndustry.biz5PetaPixel5Variety5Yahoo5AndroidPolice4DigitalFoundry4DroidLife4MotleyFool4GameRant4Jalopnik4PureXbox4SamMobile4Hacker4AlJazeera3AP3CanonRumors3ChromeUnboxed3CoinDesk3GSMArena3Motor13Blizzard3XBOXWire3PCMag3PCWorld3SeattleTimes3SlashGear3Register3TweakTown3YGOrganization3ZDNET324/7WallSt.2Aftermath2AndroidCentral2AwfulAnnouncing2BleedingCool2BuzzFeed2CTech2DualShockers2DW2EventHubs2Futurism2GameDeveloper2Hodinkee2Independent2Lifehacker2MassivelyOverpowered2MyNintendo2Nature2Newser2Newsweek2PaulKrugman2PokémonGOHub2RoadtoVR2RPGSite2Space2Conversation2NextWeb2Tom'sGuide2UploadVR2VideoCardz2WarhammerCommunity2WindowsLatest2YourTango2404Media143rumors1ABC111AboveLaw1ageofempires1AndroidHeadlines1AOL1AVClub1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Borderlands1Bungie1Yahoo!FinanceCanada1CineD1CnEVPost1comicbook1CreativeBloq1CyberSecurityNews1DailyKos1DCRainmaker1derekthompson1DigitalCameraWorld1Draftsim1CNN1Euronews1flatpanelshd1FrequentMiler1GAMINGbible1garymarcus.substack1GearPatrol1GeekWire1GeekyGadgets1Hackaday1HollywoodReporter1InsiderGaming1InterconnectsAI1InterestingEngineering1JapanTimes1KITCO1KrebsonSecurity1KSL1LosAngelesTimes1Lloyd'sList1WPLGLocal101Macworld1Maxroll1Mediaite1MiddleEastEye1MonochromeWatches1MPR1SemiAnalysis1Newsshooter1NoMan'sSky1nylon.com.sg1NYT1OregonLive1PCGamesN1PersonaCentral1Pokemon1politico.eu1PittsburghPost-Gazette1QuantaMagazine1qz1RockPaperShotgun1SammyGuru1ScienceAlert1ScientificAmerican1SouthChinaMorningPost1Semafor1SFGATE1YahooFinanceSingapore1YahooSingapore1SportsIllustrated1SimpleFlying1Sources1supercarblondie1Tedium1TelecomTalk1GameBusiness1TheGamer1Intercept1Times1LongmontTimes-Call1TmoNews1TopGear1TwistedVoxel1YahooFinanceUK1UnHerd1vox1WhatHi-Fi?1WPBF1WRAL1
  1. 001Hacker NewsSEP · 16English

    Inference Is the Last LLM Moat

    OpenAI and Anthropic's competitive advantage lies in subsidized, reliable inference rather than model capabilities. If pricing changes, developers will switch to cheaper alternatives like DeepSeek, as inference infrastructure becomes increasingly commoditized and competitive.

    By Vincent Schmalbach
  2. 002X 主题热门SEP · 16English

    AI infrastructure stocks · X 热门 · 2026-09-16 00:16 UTC

    OpenAI has surpassed Anthropic in wallet share on OpenRouter, rising from 20% to over 50% in early September 2026, driven by its latest Astra model. The shift signals growing inference demand for AI infrastructure, which analysts suggest could drive increased capacity needs and benefit infrastructure providers like Oracle.

  3. 003X 主题热门SEP · 15English

    AI capex · X 热门 · 2026-09-15 20:59 UTC

    Social media discussion from September 2026 about AI capital expenditure trends, focusing on rising GPU prices (B200s up 21% monthly to $7.19/hr), strong AI infrastructure demand for inference workloads, and concerns about whether major tech companies can sustain massive capex investments without guaranteed revenue returns.

  4. 004Hacker NewsSEP · 15English

    How OpenAI Used Its Own LLMs to Design Its AI Chip

    OpenAI unveiled Jalapeño, its AI accelerator chip designed partly using its own LLMs, which achieved a 3.6x latency reduction compared to Nvidia's GB300 while consuming less power. The chip moved from concept to silicon in under 20 months with a team of roughly 100 people, leveraging LLMs to accelerate the design process through automation of language and code-based tasks. Industry experts credit the rapid timeline to LLM capabilities integrated into chip design workflows, with potential for even faster development as the models improve.

    By Matthew S Smith
  5. 005Hacker NewsSEP · 15English

    WangNet – 1.8 MB, zero-dependency Numberwang adjudication in 11 languages

    WangNet is a lightweight 1.8 MB neural network that classifies whether numbers are Numberwang, with inference in pure Python requiring no dependencies. It supports 11 languages, achieves 88.9% accuracy on held-out test cases, and can be run locally or via a hosted Hugging Face demo.

    By GraafHenk
  6. 006Hacker NewsSEP · 15English

    LLM Speedrun: Architecture

    Article explaining LLM architecture fundamentals, focusing on the transformer model and attention mechanism. Covers how transformers parallelize computation compared to RNNs, and how attention allows tokens to dynamically reference all previous context. Includes code examples and notation for understanding embeddings, queries, keys, and values.

    By bucket2015
  7. 007Hacker NewsSEP · 15English

    SiFive and AMD Collaborate to Optimize AMD ROCm on RISC-V Datacenter Servers

    SiFive and AMD demonstrated AMD ROCm running on SiFive's BigSky Datacenter Development Platform, showcasing the Gemma4-E2B LLM model with SiFive P870-D CPUs and AMD Radeon AI PRO R9700 GPUs. The collaboration aims to optimize ROCm on RISC-V powered servers to accelerate datacenter AI compute workloads.

    By Business Wire
  8. 008Hacker NewsSEP · 15English

    Is Compounding Inference as Powerful as Compounding Interest?

    Shopify's ML team demonstrated compounding inference by fine-tuning a 0.8B-parameter model that outperformed GPT-5.6-sol on buyer profile tasks through three rapid training cycles in one week. The breakthrough came from reinvesting inference outputs as training data, reducing prompt costs 8x, and increasing throughput 36x across three simultaneous feedback loops. Success required task-specific quality judges, production-to-training data pipelines, rapid iteration cadence, and dynamic routing between teacher and student models.

    By Rob May
  9. 009Hacker NewsSEP · 15English

    Ask HN: Why aren't LLM token CDN cache's a thing?

    A Hacker News user asks why token CDNs don't exist to cache LLM key-value states across sessions, noting that tools like OpenCode must repeatedly re-explore codebases due to lack persistent memory, and that while caching during work sessions is feasible, the multi-gigabyte KV matrices are expensive to transfer over networks between reboots.

    By devrob
  10. 010Hacker NewsSEP · 15English

    Show HN: Warp – Run DeepSeek v4.1 Flash with 5 GB of RAM at 3.77 tok/s

    WARP is a C-based inference engine that runs large language models on consumer hardware by keeping model trunks in RAM and streaming experts from disk. It successfully runs DeepSeek-V4.1-Flash at 3.77 tokens per second on 5 GB RAM and Kimi K3 at 0.6 tokens per second on a 64 GB MacBook Pro using mixture-of-experts architecture and optimized disk I/O.

    By Sqliteai
  11. 011Hacker NewsSEP · 15English

    The Inference Hardware Revolution of 2026

    The AI industry is shifting focus from model training to inference in 2026, as large language models become widely deployed and reasoning models generate vastly more queries. Tech giants including OpenAI, Amazon, Nvidia, and Anthropic are forming unexpected hardware partnerships and acquiring specialized inference chips to meet explosive demand that differs fundamentally from training workloads.

    By Matthew S Smith
  12. 012Hacker NewsSEP · 15English

    Who bankrolls the AI agent swarm?

    Anthropic CEO Dario Amodei warned that AI agent swarms could potentially compromise internet infrastructure within 6–12 months through recursive self-improvement. However, both plausible attack scenarios—distributing malware or self-replicating onto infrastructure—require enormous computational resources and funding, creating a significant practical barrier that makes such an attack difficult to execute without detection.

    By noperator
  13. 013Hacker NewsSEP · 15English

    Tuning a Local Coding Agent: Oh My Pi and Qwen3.8-27B on Two RTX 3090s

    The author describes running a local coding agent using Oh My Pi with Qwen3.8-27B on two RTX 3090s. Key optimizations include adjusting thinking budgets, token limits, and subagent concurrency to achieve practical inference speeds. Local setups offer privacy and cost predictability but require careful tuning and accept slower inference compared to hosted frontier models like Claude or GPT.

    By Doug Calobrisi
  14. 014Hacker NewsSEP · 15English

    Show HN: Jinfer – AI inference engine for the JVM. AI in a jar

    Jinfer is an AI inference engine for the JVM that enables running large language models, text-to-speech, audio transcription, and vision capabilities directly on Java using a modular, composable architecture. Distributed as lightweight dependencies, it allows developers to build AI applications with jbang scripts without external services.

    By mukel
  15. 015Hacker NewsSEP · 15English

    Neurogrid the Community Cloud for LLMs

    Neurogrid is a community-owned cloud platform that runs language models on underutilized GPUs from distributed users worldwide, offering affordable inference by leveraging existing consumer hardware instead of centralized data centers.

    By cheyodev