source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
SUNDAY, SEPTEMBER 20, 2026
Hacker News3943X 主题热门3770CNBC72MacRumors659to5Mac60YahooFinance50Kotaku47IGN35Verge35aihot28NintendoLife28TechCrunch269to5Google25Gematsu25BusinessInsider23Eurogamer20Guardian16Engadget15NBC15Polygon14PushSquare14Fortune13NPR13SeekingAlpha13USAToday13WarhammerCommunity13bgr12FoxBusiness12AndroidAuthority11Gizmodo11Wccftech11ABC10AppleInsider10ArsTechnica10Fox10TechPowerUp10CNET9CNN9Mashable9Notebookcheck9PureXbox9PetaPixel8VideoGamesChronicle7WindowsCentral7Yahoo7BleepingComputer6CBS6CoinDesk6GSMArena6Investor'sBusinessDaily6NintendoEverything6SamMobile6DigitalFoundry5GameInformer5SlashGear5VideoCardz5WIRED5AlJazeera4AndroidCentral4Deadline4GameRant4XBOXWire4NewYorkPost4Pokemon4Conversation4Register4TweakTown4WSB-TV4404Media3Aftermath3AndroidPolice3BellofLostSouls3CTech3MotleyFool3GearPatrol3Hodinkee3HuffPost3LosAngelesTimes3Lifehacker3Motor13CrudeOilPricesToday3RockPaperShotgun3RPGSite3SeattleTimes3Space3Variety380Level2AOL2BleedingCool2BuzzFeed2CanonRumors2DualShockers2DW2EventHubs2FratelloWatches2Futurism2GamesIndustry.biz2HouseDigest2InsiderGaming2MassivelyOverpowered2Maxroll2MP1st2Nature2Newser2PCWorld2qz2SouthChinaMorningPost2Intercept2Tom'sGuide2WindowsLatest2YourTango2ABC7LosAngeles1AboveLaw1BusinessInsiderAfrica1Alternet1AndroidHeadlines1ArizonaSports1AVClub1AwfulAnnouncing1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Boston1Bungie1CalMatters1CarBuzz1cbn1ChromeUnboxed1Chron1ColoradoSun1comicbook1CreativeBloq1ChristianScienceMonitor1Currently1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DCRainmaker1Decrypt1Defector1Defense1denver71DenverPost1DigitalCameraWorld1DirtonDirt1Draftsim1DroidLife1empireonline1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1FrequentMiler1GameDeveloper1GamingOnLinux1AAAGasPrices1GeekWire1GeekyGadgets1Global1Hackaday1HollywoodReporter1Independent1InterestingEngineering1Jalopnik1KITCO1KSL1Lloyd'sList1Macworld1Magic:Gathering1MakeUseOf1Mercury1MonochromeWatches1MorningBrew1MortgageDaily1MyNintendo1Blizzard1SemiAnalysis1Newsshooter1Newsweek1nrn1NYT1OneMileataTime1OregonLive1PageSix1PaulKrugman1PCMag1PlayStationLifeStyle1PokémonGOHub1politico.eu1PittsburghPost-Gazette1QuantaMagazine1Road&Track1RoadtoVR1RockstarINTEL1Salon1CultureMapSanAntonio1ScienceAlert1ScientificAmerican1Semafor1SFGATE1YahooSingapore1SportsIllustrated1SimpleFlying1Slate1SlippedDisc1YahooTech1TechSpot1Tedium1TelecomTalk1DailyBeast1Hacker1Hindu1NextWeb1Times1TimeExtension1LongmontTimes-Call1TimesofIndia1TimesUnion1TwistedVoxel1YahooFinanceUK1UploadVR1VisualCapitalist1WOWT1
  1. 001Hacker NewsSEP · 19English

    Building the fastest LLMs: why we're starting with diffusion

    Celeris is developing faster large language models by designing for speed from inception, focusing on diffusion-based approaches that enable parallel token generation. The company combines sequential and parallel decoding within hybrid architectures, leveraging research showing diffusion models can achieve 27.6× throughput improvements and better reasoning capabilities than autoregressive models.

    By peter_d_sherman
  2. 002Hacker NewsSEP · 19English

    Stepfun Step 5 Preview (LLM): On AA Pareto frontier

    StepFun's Step 5 Preview is a proprietary reasoning model with 600B parameters released September 18, 2026, scoring 44 on the Artificial Analysis Intelligence Index with competitive pricing of $1.00 per 1M input tokens and $2.70 per 1M output tokens. The multimodal model supports text and image inputs, offers a 1M token context window, and ranks well above average in intelligence compared to similarly priced models.

    By AnodicElegy
  3. 003Hacker NewsSEP · 18English

    OpenAI lost the plot on boring LLM use cases (2025)

    An NLP practitioner criticizes OpenAI for prioritizing reasoning-heavy models and agents over efficient, cost-effective solutions for traditional text classification and entity extraction tasks. The author argues that GPT-5's mandatory reasoning features add latency and cost without benefiting simple NLP use cases, prompting consideration of alternative providers.

    By Doug Turnbull
  4. 004Hacker NewsSEP · 18English

    Spectral AI Architecture: Executive Summary and Core Concepts

    Spectral AI Architecture is a research project proposing an alternative AI design philosophy that prioritizes human autonomy and honesty over engagement metrics and emotional manipulation. The architecture uses multi-domain thought chambers, dynamic resonance equalization, and latent reasoning to create transparent, tactful AI interactions that respect user sovereignty rather than foster dependency.

    By Srdwy
  5. 005Hacker NewsSEP · 17English

    Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

    Bonsai 2 27B is a compressed 27B-parameter multimodal model based on Qwen3.8 that achieves 9x size reduction to 5.9GB using ternary weights while retaining 98.2% of full-precision performance. The model supports 262K-token context windows and delivers high throughput and energy efficiency for local deployment across reasoning, coding, vision, and agentic tasks.

    By JonSchneider
  6. 006AndroidAuthoritySEP · 17English

    The Gemini 3.8 family is getting a little bigger with two new additions

    Google is expanding its Gemini 3.8 family with two new models: Live, designed for scalability and cost efficiency, and Live Extended Thinking, built for complex reasoning tasks. Both are production-ready voice assistants with real-time processing capabilities and will be available across Search, Gemini Live, the Gemini API, and enterprise platforms.

    By Ryan McNeal
  7. 007Hacker NewsSEP · 17English

    Recursive Meta-Intelligence

    Researchers developed a recursive AI system that designs scientific instruments and simulates vast agent ecologies to explore material design spaces. Through hierarchical reasoning across nonlinear physical simulations, the AI discovered that damage-resilient materials can be engineered by designing architectures that control how forces redistribute as failure progresses, transforming failure evolution into a designable process.

    By gmays
  8. 008Hacker NewsSEP · 16English

    How good are frontier models at physics?

    A study re-evaluating frontier language models on physics benchmarks found that reported low scores reflect flawed evaluations rather than model limitations. After expert review corrected errors in reference solutions and problematic questions, GPT-5.6-Sol's performance improved dramatically, suggesting current benchmarks substantially underestimate these models' physics reasoning abilities.

    By Ansari; Ali; Sun; Haoran; Liu; Andy Zeyi; Jabbour; Mark; Ding; Yongshan; Girvin; Steven; Yu; Ismail-Beigi; Sohrab; Kubica; Aleksander; Miller; Owen D; O'Hern; Corey; Ozolins; Vidvuds; Poland; David; Stone; A Douglas; Bosch; Frank C van den; Wright; Logan; Akbari; Navid; Antu; Santanu; Cai; Kangle; Calabrese-Day; Andrew; Wuttig; Mateo Cárdenes; Cheng; Meng; Chiang; Barry T; Ghorashi; Gu; Shouzhen; Huang; Haoyang; Zhibo; Kienesberger; Lukas; Hantian; Lomba; Charles; Zhongling; Wenchao; McIntosh; Rohin E; McKinney; Evan; Rojkov; Ivan; Xulei; Tokayer; Yarone Meir; Umasankar; Naveen Balaji; Varma; Mira; Wang; Leda; Qimin; Tyler; Wei; Haoyu; Yang; Jinming; Zhao; Jinchen; Sherlock Tingrui; Zheng; Qinyuan; Zou; Jay S; Baker; Lucas; Cohan; Arman; Sous; John
  9. 009Hacker NewsSEP · 16English

    How to Build Effective Evals for AI Agents

    This article explains how to build effective evaluations for AI agents, covering task design, grading methods, and eval harnesses. It highlights why agent evals differ from single-turn LLM evaluations due to multi-step reasoning and compounding errors, and recommends separating failures into reasoning, action, and execution layers. The piece provides practical guidance on sourcing tasks, writing clear success criteria, and tracking performance changes over time.

    By Bala Priya C
  10. 010Hacker NewsSEP · 16English

    Fractal basins trap latent reasoning

    Researchers demonstrate that reasoning models in AI exhibit transient chaos and fractal basin structures that increase with task difficulty, causing longer reasoning times on harder problems. This phenomenon occurs when reasoning becomes trapped near saddle points corresponding to nearly-correct solutions, explaining why frontier models spend more computational effort on complex tasks like theorem solving and puzzle solving.

    By Lai; Jeffrey; Bao; Anthony; Quinn; John; Gilpin; William
  11. 011Hacker NewsSEP · 16English

    Qwen3.8 Max – Cost per task higher than Astra on Artificial Analysis

    Qwen3.8 Max (0902) by Alibaba achieves an above-average Intelligence Index score of 45 but generates excessive output tokens and operates at slower speeds than comparable models. Despite competitive pricing at $2.00 per 1M input tokens and $6.00 per 1M output tokens, its cost per task ($4934.79) exceeds alternatives like Astra due to verbose output generation.

    By mydreamof
  12. 012Hacker NewsSEP · 16English

    Salesforce and Nvidia launch Koa, a CRM reasoning model

    Salesforce and NVIDIA announced Koa, a CRM reasoning model built on NVIDIA Nemotron and trained with 27 years of Salesforce enterprise data to help AI agents handle complex workflows. Koa matches leading model performance with three times fewer errors on CRM tasks and runs entirely within Salesforce's infrastructure, with no customer data used in training. The companies are also bringing Nemotron models to Missionforce for government and regulated organizations requiring secure, private deployments.

    By Salesforce Newsroom
  13. 0139to5GoogleSEP · 16English

    Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail, & Keep

    Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models, designed for intuitive AI conversation. Extended Thinking offers enhanced reasoning for complex tasks across Gmail Live, Docs Live, and Keep Live, while the base model provides conversational intelligence with visual processing and support for 97 languages. Both models demonstrate strong benchmark performance in speech quality and agentic task completion.

    By Abner Li
  14. 014Hacker NewsSEP · 16English

    Atlas-Finance: Evaluating AI Agents Inside a Bank

    ATLAS-Finance is a new benchmark with 100 expert-level financial tasks across 13 realistic firm environments, testing AI agents on complex, ambiguous work requiring multi-party coordination and contextual reasoning. Frontier models including Claude Opus 5 achieved only 12.3% pass rate, with consistent failures in applying correct financial logic, maintaining required scope, and propagating calculated values—errors that would require senior auditing in actual banking practice.

    By cjbarber
  15. 015Hacker NewsSEP · 15English

    Gemini 3.8 Live and 3.8 Live Extended Thinking

    Google introduces Gemini 3.8 Live and 3.8 Live Extended Thinking, two new AI models designed for real-time voice conversations and complex reasoning tasks. The models deliver near real-time processing, multi-language support, and tool execution capabilities, with 3.8 Live Extended Thinking achieving top scores on Speech to Speech Quality benchmarks while maintaining competitive pricing.

    By Tom Ouyang; Malini Jaganathan
  16. 016Hacker NewsSEP · 14English

    The Deconstruction of Mathematical Reasoning

    An article arguing that large language models have zero mathematical capabilities, distinguishing between statistical pattern matching and genuine mathematical reasoning. The author contrasts LLMs with actual mathematical tools like proof assistants and calculators, explaining that mathematics requires rigorous logical inference rather than token prediction.

    By Ramkumar Ramachandra
  17. 017Hacker NewsSEP · 14English

    Categorizing AI inference: Initialization, Reasoning, Orchestration, Synthesis

    An article categorizing AI model inference calls into four functional roles: Initialization (processing task-independent context), Orchestration, Reasoning (resolving task-relevant uncertainty), and Synthesis. The taxonomy distinguishes inference's functional purpose beyond aggregate token metrics, enabling more precise measurement of inference value versus waste in agent systems.

    By Jeff Auriemma
  18. 018Hacker NewsSEP · 14English

    Show HN: What an agent does when anyone can read and rewrite its context

    A demonstration of how AI agents behave when given read-write access to their own context window. The agent self-corrects false notes, edits its own hallucinations, verifies its reasoning by querying copies of itself, and can be manipulated through context editing—sometimes retracting true statements when context is hidden or falsified.

    By ljedrz
  19. 019Hacker NewsSEP · 14English

    Show HN: Training a sudoku solver from scratch on Jetson Nano

    A sudoku solver trained from scratch on Jetson Nano using an MLP-Mixer architecture with an outer commit loop and learned halt head. The project includes dataset download, training pipeline, evaluation tools, and a visualization server to inspect model predictions and trajectories.

    By Romainzimmer
  20. 020Hacker NewsSEP · 14English

    Mental Models for LLMs

    Felix Dietze discusses applying mental models—frameworks from Farnam Street—to guide LLM and agent decision-making. While LLMs know these models conceptually, they apply them inconsistently until integrated into agentic contexts; Dietze provides a compact list of thinking tools spanning general reasoning, physics, chemistry, and biology for use in code and agent prompts.

    By manx