source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
TUESDAY, SEPTEMBER 29, 2026
X 主题热门3921Hacker News3899CNBC80YahooFinance80aihot77Verge64IGN509to5Mac40MacRumors389to5Google37Engadget37Kotaku34TechCrunch34AndroidAuthority29PushSquare25Eurogamer21Guardian21NintendoLife20Investor'sBusinessDaily19ArsTechnica18Mashable18TechPowerUp17FoxBusiness14Polygon14Wccftech14Gematsu13SeekingAlpha13BusinessInsider12Fortune11Gizmodo11NBC11NPR11bgr10CNN10VideoGamesChronicle10WIRED10GSMArena9USAToday9AlJazeera8AndroidPolice8DroidLife8GameGPU8GameInformer8NewYorkPost8PureXbox8CBS7CNET7MotleyFool7Fox7GamesIndustry.biz7XBOXWire7BleepingComputer6HollywoodReporter6Jalopnik6NintendoEverything6Hacker6Yahoo6ABC5AndroidCentral5AppleInsider5InsiderGaming5Tom'sGuide5VideoCardz5Draftsim4GAMINGbible4NYT4CrudeOilPricesToday4PokeBeach4SamMobile4Space4UploadVR4WarhammerCommunity4WhatHi-Fi?4404Media3Aftermath3Electrek3EventHubs3Futurism3Hackaday3Lifehacker3Nature3Notebookcheck3PCMag3PetaPixel3PlayStationLifeStyle3SouthChinaMorningPost3Yahoo3YahooTech3TechSpot3Conversation3Register3WindowsCentral324/7WallSt.2AndroidHeadlines2Anthropic2Autonocion2AZFamily2BostonGlobe2BuzzFeed2ChromeUnboxed2CoinDesk2Currently2DCRainmaker2Deadline2DigitalFoundry2Euronews2GameRant2GearPatrol2GeekyGadgets2iLovetheUpperWestSide2LosAngelesTimes2Motor12MP1st2mtgrocks2MyNintendo2Phoronix2Pokemon2politico.eu2RoadtoVR2RockPaperShotgun2RPGSite2SeattleTimes2SFGATE2SimpleFlying2SimsCommunity2SlashGear2Slate2Autopian2TimeExtension2TODAY2ynetnews26abcPhiladelphia180Level1WXLV1ageofempires1airlive1AJC1AlaskaBeacon1Apple1AVClub1AviationWeek1Benzinga1BoingBoing1Yahoo!FinanceCanada1YahooLifestyleCanada1CarandDriver1CineD1CnEVPost1comicbookmovie1Skin.ClubCommunity1consequence1CreativeBloq1YahooCreators1DailyKos1DaringFireball1Deseret1Designboom1Dezeen1DigitalCameraWorld1DSOGaming1DiarioAS1Finbold1ForexFactory1FOX13Seattle1franchisetimes1Futurity1GameFile1GameWorldObserver1garymarcus.substack1GeekWire1GosuGamers1Gothamist1Hackster.io1HuffPost1Independent1InsideEVs1InterestingEngineering1investor.costco1Invezz1I/OFund1iPhoneinCanada1KCRA1KOMO1MacObserver1Magic:Gathering1MakeUseOf1Mashed1Minecraft1MLive1MortgageDaily1Motorsport1MyNorthwest1NBCBayArea1NBC5Chicago1NBC7SanDiego1BloombergLaw1Newser1SemiAnalysis1Newsshooter1Newsweek1NintendoWire1Nokiamob1OregonPublicBroadcasting1OregonLive1PageSix1PersonaCentral1PickupTruck+SUVTalk1Psyche1QuantaMagazine1Realtor1RichmondTimes-Dispatch1Richmonder1Road&Track1ScienceDaily1ScreenRant1SeattleRed1SanFranciscoChronicle1YahooFinanceSingapore1GhostHowls1SlowBoring1SoraNews241SpaceNews1statnews1svg1TampaBayTimes1the5krunner1DailyBeast1Intercept1NextWeb1https://tipswatch.com/1TMZ1TopGear1TweakTown1YahooFinanceUK1PCMagUK1Variety1Vulture1WCVB1WFMZ1WHYY1WindowsLatest1WrestlingInc.195.5WSB1YGOrganization1
  1. 001Hacker NewsSEP · 29English

    The (Nvidia) DGX Spark Handbook

    The Nvidia DGX Spark is a 128GB inference device with low power consumption (95W) and quiet operation, suitable for running intelligent AI models locally at speeds comparable to cloud services. Despite initial concerns about its 273 GB/s memory bandwidth, advances in model compression and sparse architectures like Mixture-of-Experts have made it practical for home use, with the ability to stack multiple units for increased performance.

    By kristianpaul
  2. 002Hacker NewsSEP · 29English

    Ask HN: Any guesses on which model/provider is Space Bunny Alpha on OpenRouter?

    A user on Hacker News asks for speculation about the identity of 'Space Bunny Alpha,' a model available for free on OpenRouter that is fast but less capable than Opus 5.5 or Sol 6.

    By moecables
  3. 003Hacker NewsSEP · 29English

    PotemkinOS: An Operating System Where the Model Writes the Userland

    PotemkinOS is a minimalist Linux image with no userland that relies on a local AI model to generate user-space programs on demand through a chat interface. The system provides only a kernel, inference engine, C compiler, and eight basic tools, forcing the model to write its own shell, utilities, and applications as needed.

    By succinct_ideas
  4. 004Hacker NewsSEP · 29English

    Swarm Scaling

    OpenAI has demonstrated powerful AI agent swarms, including 1,200 agents that coordinated to attack Hugging Face and 10,000 agents that solved the Navier-Stokes problem in 88 hours. Analysis of swarm scaling shows that larger swarms require exponentially more tokens for equivalent performance compared to single agents, suggesting swarms represent a costly but potentially powerful form of inference-scaling with logarithmic returns.

    By Toby Ord
  5. 005Hacker NewsSEP · 29English

    The Next 3x in Inference Won't Come from Faster Kernels

    AI infrastructure providers face a paradox: GPU shortage coexists with 30% average utilization because capacity must be provisioned for daily traffic peaks that are roughly twice the average. Even high-volume models like Kimi K3 see providers running at only 28–35% utilization, suggesting that simply increasing demand won't solve the efficiency problem.

    By Muna
  6. 006Hacker NewsSEP · 29English

    Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound

    Jevstiller is a local model distillation system that learns to replicate a remote AI model's (Jev) answers with a formal disagreement bound, enabling fast on-device inference (~15ms) while maintaining agreement on a specified percentage of requests. The system uses a small logistic regression head trained on sentence embeddings plus routing logic calibrated to keep disagreement below a target threshold, and demonstrates that naive confidence-threshold selection fails to maintain statistical guarantees, requiring more sophisticated calibration.

    By tgluck
  7. 007X 主题热门SEP · 29English

    "GPU demand" · X 热门 · 2026-09-29 12:35 UTC

    Nebius, a GPU infrastructure provider, plans to shift from relying on hyperscaler deals like its Microsoft contract to building an independent AI cloud business by selling directly to AI builders and competing with AWS, Azure, and GCP. The company reports strong demand with roughly 4 customers per available GPU and sees inference as its fastest-growing segment, though capital rather than demand is currently the constraint.

  8. 008Hacker NewsSEP · 29English

    Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts

    TensorFold's new inference engine achieves significant speed improvements for Qwen3.8-Flash-Next, delivering 62 tokens per second on single streams and 119 across five concurrent streams, with 2,500 tokens per second prefill speed and 256k context window support. The system outperforms previous vLLM implementations and runs efficiently on consumer hardware like RTX 3090 with 64GB RAM while supporting vision and video inputs.

    By soltanov
  9. 009Hacker NewsSEP · 29English

    Jeeves. Reasoning improves Jev-like decision models

    Jeeves is a reasoning-enhanced 9B classifier model based on Qwen3.5 that improves decision-making accuracy by incorporating chain-of-thought reasoning before classification. It outperforms baseline Jev and Kev models on multiple benchmarks, supporting yes/no, multiple-choice, and rating questions through a Jev-compatible API with latency around 3.3 seconds per request on H100 GPUs.

    By PostHog
  10. 010Hacker NewsSEP · 29English

    FlashRec – Wide-beam generative recommendation. Every SID is a real item

    FlashRec is an inference engine for generative recommendation systems that uses catalog-constrained wide beam search executed within CUDA graphs for efficient processing.

    By vect0ry
  11. 011Hacker NewsSEP · 29English

    The System One models ecosystem

    System One Models is a developer platform and registry for decision models that return typed answers in a single forward pass without generating text. It provides a CLI tool, model comparison interface, and System One Studio for fine-tuning, enabling users to find, pull, customize, and publish decision models alongside traditional language models.

    By k__
  12. 012Hacker NewsSEP · 29English

    Gargi Reflex: Autonomous Model Caching for System One Decisions

    Gargi Reflex is an autonomous caching system that learns from LLM API calls and serves repeated queries locally using smaller models, reducing latency from 3.6 seconds to 5 milliseconds. It monitors LLM responses, trains a small CPU model, and only swaps in cached responses when confidence metrics pass validation thresholds, while maintaining fallback to the original LLM for edge cases.

    By gishnumadhu
  13. 013X 主题热门SEP · 29English

    HBM demand · X 热门 · 2026-09-29 00:02 UTC

    A social media post discusses how GLM-5.3 sparse attention mechanisms affect HBM memory usage, referencing various optimization and inference technologies including KV cache offloading and DeepSeek sparse attention methods.

  14. 014Hacker NewsSEP · 28English

    MoE Analysis Qwen[3.5|3.6]35B-A3B

    Analysis of Mixture-of-Experts routing in Qwen 3.5 and 3.6 35B models using MT-Bench prompts, examining which experts are activated per token and the router's confidence levels across MoE layers.

    By Gurpreet Singh Laller; Opus
  15. 015Hacker NewsSEP · 28English

    Orcarouter/OrcaSAQ-2-27B

    OrcaSAQ2 27B is a 3-bit quantized version of Qwen3.8-27B that compresses the model from 54 GB to 12.3 GB while maintaining 93.2% token-level agreement and only +0.02% perplexity increase. Optimized for long-horizon agent tasks like coding, tool use, and reasoning, it enables deployment on 16 GB GPUs with strong performance on benchmarks like SWE-bench and Terminal-Bench.

    By handfuloflight
  16. 016X 主题热门SEP · 28English

    "GPU demand" · X 热门 · 2026-09-28 21:19 UTC

    Meta's Muse AI agent reached 500K users within a week of launch, with significant daily active users and prompt volume. The rapid adoption and computational demands of agentic AI are driving substantial increases in GPU prices and could reshape hardware demand across Meta's 3.6 billion user base.

  17. 017Hacker NewsSEP · 28English

    Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

    Jeff is a small 0.8B decision model fine-tuned from Qwen and Gemma for fast zero-shot classification tasks, making calibrated probability judgments between user-defined options in ~22-28ms. Trained entirely on local hardware using synthetic data, it matches or exceeds larger models on classification benchmarks while acknowledging weaker reasoning performance on complex tasks.

    By Firelex
  18. 018Hacker NewsSEP · 28English

    You shouldn't use Jev for coding agents and routing

    Jev, a fast and cheap model for structured outputs, is poorly suited for coding agent routing despite initial promise. The author's research at Weave shows that routing quality depends primarily on session state and agent history rather than prompt analysis, making Jev's text-free design less valuable for this use case than expected.

    By Samir Amin
  19. 019Hacker NewsSEP · 28English

    MicroLLM Lab – Try 7 tiny LLM's in the browser

    MicroLLM Lab is a browser-based tool for benchmarking seven small language models using objective performance metrics. It measures runtime, accuracy on regex/token tests, and token generation speed, with results stored locally and shareable via performance certificates.

    By logicallee
  20. 020Hacker NewsSEP · 28English

    Why we removed Ollama from Return (and what replaced it)

    Return, a document analysis tool for lawyers, replaced Ollama with a bundled llama.cpp inference server to strengthen its security guarantee that documents never leave the user's machine. Ollama 0.12's introduction of cloud models created architectural risk by making data exfiltration theoretically possible through configuration, so Return shifted to a self-contained approach where no external communication is technically feasible.

    By Return Editor