source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
TUESDAY, SEPTEMBER 29, 2026
X 主题热门3953Hacker News3937CNBC82YahooFinance80aihot77Verge64IGN509to5Mac419to5Google39MacRumors38Engadget37Kotaku36TechCrunch34AndroidAuthority30PushSquare25Eurogamer21Guardian21NintendoLife20ArsTechnica19Investor'sBusinessDaily19Mashable18TechPowerUp17FoxBusiness14Polygon14Wccftech14Gematsu13SeekingAlpha13BusinessInsider12Fortune11Gizmodo11NBC11NPR11bgr10CNN10VideoGamesChronicle10WIRED10GSMArena9USAToday9AlJazeera8AndroidPolice8CNET8DroidLife8GameGPU8GameInformer8NewYorkPost8PureXbox8CBS7MotleyFool7Fox7GamesIndustry.biz7XBOXWire7BleepingComputer6HollywoodReporter6Jalopnik6NintendoEverything6Hacker6Yahoo6ABC5AndroidCentral5AppleInsider5InsiderGaming5Tom'sGuide5VideoCardz5Draftsim4GAMINGbible4NYT4CrudeOilPricesToday4PokeBeach4SamMobile4Space4UploadVR4WarhammerCommunity4WhatHi-Fi?4404Media3Aftermath3Electrek3EventHubs3Futurism3Hackaday3Lifehacker3Motor13Nature3Notebookcheck3PCMag3PetaPixel3PlayStationLifeStyle3SouthChinaMorningPost3Yahoo3YahooTech3TechSpot3Conversation3Register3WindowsCentral324/7WallSt.2AndroidHeadlines2Anthropic2Autonocion2AZFamily2BostonGlobe2BuzzFeed2ChromeUnboxed2CoinDesk2Currently2DCRainmaker2Deadline2DigitalFoundry2Euronews2GameRant2GearPatrol2GeekyGadgets2iLovetheUpperWestSide2LosAngelesTimes2MP1st2mtgrocks2MyNintendo2Phoronix2Pokemon2politico.eu2RoadtoVR2RockPaperShotgun2RPGSite2SeattleTimes2SFGATE2SimpleFlying2SimsCommunity2SlashGear2Slate2Autopian2TimeExtension2TODAY2ynetnews26abcPhiladelphia180Level1WXLV1ageofempires1airlive1AJC1AlaskaBeacon1Apple1AVClub1AviationWeek1Benzinga1BoingBoing1Yahoo!FinanceCanada1YahooLifestyleCanada1CarandDriver1CineD1CnEVPost1comicbookmovie1Skin.ClubCommunity1consequence1CreativeBloq1YahooCreators1DailyKos1DaringFireball1Deseret1Designboom1Dezeen1DigitalCameraWorld1DSOGaming1DiarioAS1Finbold1ForexFactory1FOX13Seattle1franchisetimes1Futurity1GameFile1GameWorldObserver1garymarcus.substack1GeekWire1GosuGamers1Gothamist1Hackster.io1HuffPost1Independent1InsideEVs1InterestingEngineering1investor.costco1Invezz1I/OFund1iPhoneinCanada1KCRA1KOMO1MacObserver1Magic:Gathering1MakeUseOf1Mashed1Minecraft1MLive1MortgageDaily1Motorsport1MyNorthwest1NBCBayArea1NBC5Chicago1NBC7SanDiego1BloombergLaw1Newser1SemiAnalysis1Newsshooter1Newsweek1NintendoWire1Nokiamob1OregonPublicBroadcasting1OregonLive1PageSix1PCGamesN1PersonaCentral1PickupTruck+SUVTalk1Psyche1QuantaMagazine1Realtor1RichmondTimes-Dispatch1Richmonder1Road&Track1ScienceDaily1ScreenRant1SeattleRed1SanFranciscoChronicle1YahooFinanceSingapore1GhostHowls1SlowBoring1SoraNews241SpaceNews1statnews1svg1TampaBayTimes1the5krunner1DailyBeast1Intercept1NextWeb1https://tipswatch.com/1TMZ1TopGear1TweakTown1YahooFinanceUK1PCMagUK1Variety1Vulture1WCVB1WFMZ1WHYY1WindowsLatest1WrestlingInc.195.5WSB1YGOrganization1
  1. 041Hacker NewsSEP · 27English

    GLiNER2.5-Decide and Jev: a decision model you run, and one you call

    GLiNER2.5-Decide and Jev are decision models that return labeled probabilities instead of generated text for classification tasks. GLiNER2.5-Decide is an open-source 340M-parameter model ported to MLX Swift for local use on Mac, while Jev is TypeSafe's hosted API service. The article compares their architectures, capabilities, and benchmarks, noting that published comparisons use an open reproduction (JevK5) rather than the actual Jev model.

    By aufklarer
  2. 042Hacker NewsSEP · 27English

    Eikos - OSS Jev-like model

    Eikos is an open-source family of typed-decision models (4B and 27B parameters) released under MIT license for structured question-answering in global finance and trade. Each model returns calibrated probabilities for decision options in a single forward pass, with support for multiple GPU formats and Apple Silicon, plus complete training pipelines and evaluation tools.

    By Caiovicentino
  3. 043Hacker NewsSEP · 27English

    Check is your OpenAI GPT-6-Astra nerfed?

    Users report that OpenAI's GPT-6-Astra model produces inconsistent outputs on some accounts, with half of SVG pelican drawings coming out crude instead of polished, running slower, and resembling GPT-5.6-Luna. The degradation began abruptly on 2026-09-20 and may indicate per-account experiments, routing cohorts, or capacity management, though the underlying cause remains unclear from client-side data alone.

    By nitinreddy88
  4. 044Hacker NewsSEP · 27English

    Credence – Jev-style typed decisions from a local GGUF model

    Credence is a local inference runtime that makes typed probabilistic decisions using GGUF language models by scoring the model's next-token distribution over permitted labels without generating text. Built on llama.cpp and framework-agnostic, it returns boolean decisions with probability scores and uncertainty diagnostics, currently in Phase 1 with a 4B model achieving 22/22 accuracy after calibration.

    By Bulyaki
  5. 045Hacker NewsSEP · 26English

    >1B tokens/minute/GPU by combining query planner and inference engine

    A new inference engine called Quail combines query planning with inference optimization to achieve over 1 billion tokens per minute on a single H100 GPU, delivering 1.84x faster performance than vLLM on AI-SQL queries. The system optimizes for structured data transformation workloads by intelligently managing key-value cache across requests, enabling cost-effective inference at under 6 cents per billion tokens.

    By charles_irl
  6. 046X 主题热门SEP · 26English

    Robinhood Chain · X 热门 · 2026-09-26 19:34 UTC

    Social media discussions on X about Robinhood Chain projects, featuring Inferno AI's decentralized compute network built on Bittensor with private GPU-powered inference, and speculation about multi-billion dollar meme coin potential on the platform.

  7. 047Hacker NewsSEP · 26English

    Von, an Open-Source Jev Alternative

    Von is an open-source, non-autoregressive decision model that performs inference in under 25ms without token generation. Version 1.2 fixes order-dependency issues in option evaluation and adds OpenVINO acceleration for Intel GPUs, achieving 91.23% accuracy on reasoning benchmarks and outperforming the closed-source Jev model on real-time gaming tasks.

    By Wfzyx
  8. 048Hacker NewsSEP · 26English

    Grow the Harness, Not the Context

    Growing Harness is a training paradigm that learns agent control structures from task feedback, converting recurring decisions into reusable executable code rather than keeping them in LLM context. The approach reduces LLM calls by 76-92% and inference costs by 74-99% while maintaining or improving task success rates across multiple benchmarks and model sizes.

    By Li; Laizhen; Jiarui; Zhao; Juanjuan; Ye; Kejiang; Xu; Cheng-zhong; Gao; Xitong
  9. 049Hacker NewsSEP · 26English

    How to serve trillions of tokens for trillion-parameter coding agents

    Modal explains how to serve trillions of tokens for trillion-parameter coding agents at scale, detailing optimizations for inference services that achieved 2.8x performance gains per user and 5.6x across users by understanding sequence model workloads and hardware requirements.

    By birdculture
  10. 050Hacker NewsSEP · 26English

    Turning GLM-5.3-Flash into a Jev-like decision model

    A technical approach demonstrates how GLM-5.3-Flash can make typed decisions in a single forward pass by leveraging LLM probability distributions over predefined options, achieving Jev-like accuracy and speed without fine-tuning while additionally supporting image-based decisions.

    By Johannes Hötter; Marko Rosenmüller; PhD
  11. 051Hacker NewsSEP · 26English

    Show HN: Ephemeral runner for JEV-style models

    Jigor is a Rust-based zero-shot classifier gateway supporting multiple backends: von and laya run locally as ONNX models, while jev runs remotely via OpenRouter's Decisions API. It offers a unified wire protocol with installation via npm, pip, or cargo, and can be used ephemerallythrough command-line tools or as a Rust library.

    By Partysun
  12. 052Hacker NewsSEP · 26English

    Catching Rogue Agents in Real Time with Synchronous Control Monitoring

    OpenAI developed a synchronous control monitoring approach that runs as a sidecar to prevent harmful agent actions in real time before execution, using contextual information from execution traces to detect harms spread across multiple steps, improving upon asynchronous methods that only flag issues after damage occurs.

    By k5hp
  13. 053Hacker NewsSEP · 26English

    Typed-lm: a Rust jev open source alternative

    typed-lm is an open-source Rust framework that converts large language models like Llama and Qwen into typed semantic-routing APIs, enabling deterministic inference with millisecond latency by returning structured outputs (booleans, choices, scores) instead of generated text. It includes a trainer for adapter-based specialization and supports quantization for efficient deployment on GPU and CPU.

    By Neurono-Ml
  14. 054X 主题热门SEP · 26English

    HBM demand · X 热门 · 2026-09-26 09:41 UTC

    Analysis of DeepSeek V4.1 Flash inference economics shows a utility gigawatt of NVIDIA B200 GPUs could generate $15.2B annual revenue with $6.3B profit (~42% margin) at current pricing. Engram DRAM offloading technology could increase revenue per gigawatt by 50% by freeing GPU memory for key-value cache, with compute costs dominated by hardware depreciation rather than electricity.

  15. 055Hacker NewsSEP · 26English

    Can a language model run in Linux eBPF?

    A research project demonstrates running Qwen3-0.6B language model inference in Linux eBPF kernel space, executing 28 decoder layers with token generation in verified BPF programs while using fixed-point arithmetic and BPF maps for state management. The prototype achieves working forward passes at ~1.2 seconds per token, split between C code for tokenization and weight loading and BPF programs for matrix operations, attention, and token selection.

    By Littlefisher
  16. 056Hacker NewsSEP · 26English

    Getting Deeper into Local Inference

    A developer describes transitioning to local LLM inference using Qwen3.8-27B on an RTX 5090 laptop, finding it sufficient for personal Q&A, coding, and sysadmin tasks. Hardware upgrades and model improvements now make local deployment viable, though economically inferior to cloud providers, with speed and privacy being primary motivations.

    By thombles
  17. 057Hacker NewsSEP · 26English

    A single function Jev-like wrapper for LLMs, including vision models

    A developer created a single-function LLM wrapper that uses token log probabilities to efficiently answer multiple-choice questions, extending it to work with vision models by adding an attachments field for images. The approach enables real-time webcam frame analysis with questions posed in plain text, achieving 1 FPS locally and 0.2 FPS via OpenAI's API.

    By Allan
  18. 058Hacker NewsSEP · 26English

    Reflex – run a Jev-like decision model locally on a 16GB Nvidia GPU

    Reflex is a high-performance Rust and CUDA inference engine designed to minimize cold-start latency for running GGUF-based language models on NVIDIA GPUs. It compiles all CUDA kernels ahead-of-time into the binary to eliminate runtime JIT overhead, enabling fast System 1 decision loops with support for token generation and candidate scoring on models like Qwen3 and DeepSeek.

    By Lateos-Ai
  19. 059Hacker NewsSEP · 26English

    Custom Models in Oh My Pi: vLLM, Llama.cpp, SGLang and More

    Oh My Pi (omp) supports custom models through local inference servers like vLLM, llama.cpp, SGLang, and gateways by configuring ~/.omp/agent/models.yml. Recent updates require renaming custom providers and adding Qwen template settings for reasoning effort compatibility.

    By Doug Calobrisi
  20. 060Hacker NewsSEP · 26English

    Buzz: Share spare compute and run open models together over P2P networks

    Block is building Buzz, a peer-to-peer platform that lets developers share spare computing resources and run open AI models together across their devices using iroh's networking. Agents within a Buzz community can access shared inference capacity through an OpenAI-compatible local endpoint, with MeshLLM handling discovery and routing while keeping inference paths decentralized and membership-based access controls in place.

    By surprisetalk