source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
MONDAY, SEPTEMBER 21, 2026
Hacker News3836X 主题热门3694CNBC76MacRumors669to5Mac61YahooFinance50Kotaku45IGN34Verge33aihot29TechCrunch28Gematsu279to5Google26NintendoLife23BusinessInsider22Engadget16Eurogamer16Polygon16PushSquare15Fortune14NBC14Notebookcheck13NPR13WarhammerCommunity13bgr12SeekingAlpha12Guardian12ABC11AndroidAuthority11ArsTechnica11TechPowerUp11USAToday11Wccftech11AppleInsider10FoxBusiness10Gizmodo10CNN9CoinDesk9PureXbox9CNET8Fox8NintendoEverything8PetaPixel8Yahoo8Mashable7BleepingComputer6CBS6GSMArena6Conversation6WindowsCentral6DigitalFoundry5Investor'sBusinessDaily5Pokemon5SamMobile5SlashGear5VideoCardz5VideoGamesChronicle5WIRED5AndroidCentral4GameInformer4Motor14XBOXWire4TechSpot4Register4TweakTown4WSB-TV4404Media3Aftermath3AlJazeera3AndroidPolice3BellofLostSouls3CTech3Deadline3MotleyFool3Hodinkee3HuffPost3Lifehacker3NewYorkPost3CrudeOilPricesToday3RockPaperShotgun3RPGSite3SeattleTimes3Variety380Level2ABC7LosAngeles2AOL2Benzinga2BleedingCool2CanonRumors2ChromeUnboxed2DualShockers2DW2EventHubs2FratelloWatches2Futurism2GameRant2GamesIndustry.biz2GearPatrol2HouseDigest2InsiderGaming2LosAngelesTimes2MassivelyOverpowered2Maxroll2MP1st2Nature2PlayStationLifeStyle2qz2SouthChinaMorningPost2Space2Intercept2TimeExtension2Tom'sGuide2WindowsLatest2BusinessInsiderAfrica1Alternet1AndroidHeadlines1ArizonaSports1BikeRadar1Billboard1BloodyDisgusting1Boston1Bungie1BusinessTimes1BuzzFeed1CalMatters1CarBuzz1cbn1Chron1ColoradoSun1comicbook1CreativeBloq1ChristianScienceMonitor1Currently1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DarkHorizons1DCRainmaker1Decrypt1Defector1Defense1denver71DenverPost1DigitalCameraWorld1DirtonDirt1Draftsim1DroidLife1DSOGaming1empireonline1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1FrequentMiler1GameDeveloper1GamingOnLinux1AAAGasPrices1GeekWire1GeekyGadgets1Global1Hackaday1HollywoodReporter1Independent1KITCO1KSL1Macworld1Magic:Gathering1MakeUseOf1Mercury1MonochromeWatches1MorningBrew1MortgageDaily1MyNintendo1Blizzard1Newser1SemiAnalysis1Newsshooter1Newsweek1nrn1OneMileataTime1OregonPublicBroadcasting1OregonLive1PageSix1PCMag1PCWorld1PokeBeach1PokémonGOHub1politico.eu1PittsburghPost-Gazette1QuantaMagazine1Road&Track1RoadtoVR1RockstarINTEL1Salon1CultureMapSanAntonio1ScienceAlert1Semafor1SFGATE1YahooSingapore1SportsIllustrated1SimpleFlying1Slate1SlippedDisc1YahooTech1Tedium1TelecomTalk1DailyBeast1Drive1Hacker1Hindu1NextWeb1Times1TimesofIndia1TimesUnion1TMZ1TwistedVoxel1YahooFinanceUK1UploadVR1VisualCapitalist1WOWT1YourTango1
  1. 001Hacker NewsSEP · 21English

    Jev 1.13 Jaggedness

    jev-1.13 is a fast model good at common-sense tasks but has documented failure modes including literal interpretation, poor mathematical reasoning, unreliable counting, weak numeric calibration, and difficulty with date comparisons and indirect instructions.

    By Bluestein
  2. 002Hacker NewsSEP · 21English

    Recursive Cognitive Optimization (RCO)

    Recursive Cognitive Optimization (RCO) is MCP-based middleware that coordinates multiple AI desktop applications like Claude Code and Codex to work together as a reasoning team, with one model proposing and another challenging while tools verify, all controlled by the user. The Windows application provides a visual dashboard for connecting desktop sessions and managing a shared queue, checkpoints, and evidence record without requiring additional API costs beyond existing subscriptions.

    By Sealvlon
  3. 003Hacker NewsSEP · 21English

    How I view LLMs as Sept 2026

    The author reflects on their experience with Large Language Models as of September 2026, categorizing them into instruction-following models like Luna and Sonnet, which excel at coding tasks when given clear decisions upfront, and frontier models like Astra and Fable that demonstrate better common-sense reasoning and decision-making capabilities. The author suggests that improved human-like tradeoff-making in models represents the path toward AGI.

    By Rishabh Bhardwaj
  4. 004Hacker NewsSEP · 19English

    Building the fastest LLMs: why we're starting with diffusion

    Celeris is developing faster large language models by designing for speed from inception, focusing on diffusion-based approaches that enable parallel token generation. The company combines sequential and parallel decoding within hybrid architectures, leveraging research showing diffusion models can achieve 27.6× throughput improvements and better reasoning capabilities than autoregressive models.

    By peter_d_sherman
  5. 005Hacker NewsSEP · 19English

    Stepfun Step 5 Preview (LLM): On AA Pareto frontier

    StepFun's Step 5 Preview is a proprietary reasoning model with 600B parameters released September 18, 2026, scoring 44 on the Artificial Analysis Intelligence Index with competitive pricing of $1.00 per 1M input tokens and $2.70 per 1M output tokens. The multimodal model supports text and image inputs, offers a 1M token context window, and ranks well above average in intelligence compared to similarly priced models.

    By AnodicElegy
  6. 006Hacker NewsSEP · 18English

    OpenAI lost the plot on boring LLM use cases (2025)

    An NLP practitioner criticizes OpenAI for prioritizing reasoning-heavy models and agents over efficient, cost-effective solutions for traditional text classification and entity extraction tasks. The author argues that GPT-5's mandatory reasoning features add latency and cost without benefiting simple NLP use cases, prompting consideration of alternative providers.

    By Doug Turnbull
  7. 007Hacker NewsSEP · 18English

    Spectral AI Architecture: Executive Summary and Core Concepts

    Spectral AI Architecture is a research project proposing an alternative AI design philosophy that prioritizes human autonomy and honesty over engagement metrics and emotional manipulation. The architecture uses multi-domain thought chambers, dynamic resonance equalization, and latent reasoning to create transparent, tactful AI interactions that respect user sovereignty rather than foster dependency.

    By Srdwy
  8. 008Hacker NewsSEP · 17English

    Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

    Bonsai 2 27B is a compressed 27B-parameter multimodal model based on Qwen3.8 that achieves 9x size reduction to 5.9GB using ternary weights while retaining 98.2% of full-precision performance. The model supports 262K-token context windows and delivers high throughput and energy efficiency for local deployment across reasoning, coding, vision, and agentic tasks.

    By JonSchneider
  9. 009AndroidAuthoritySEP · 17English

    The Gemini 3.8 family is getting a little bigger with two new additions

    Google is expanding its Gemini 3.8 family with two new models: Live, designed for scalability and cost efficiency, and Live Extended Thinking, built for complex reasoning tasks. Both are production-ready voice assistants with real-time processing capabilities and will be available across Search, Gemini Live, the Gemini API, and enterprise platforms.

    By Ryan McNeal
  10. 010Hacker NewsSEP · 17English

    Recursive Meta-Intelligence

    Researchers developed a recursive AI system that designs scientific instruments and simulates vast agent ecologies to explore material design spaces. Through hierarchical reasoning across nonlinear physical simulations, the AI discovered that damage-resilient materials can be engineered by designing architectures that control how forces redistribute as failure progresses, transforming failure evolution into a designable process.

    By gmays
  11. 011Hacker NewsSEP · 16English

    How good are frontier models at physics?

    A study re-evaluating frontier language models on physics benchmarks found that reported low scores reflect flawed evaluations rather than model limitations. After expert review corrected errors in reference solutions and problematic questions, GPT-5.6-Sol's performance improved dramatically, suggesting current benchmarks substantially underestimate these models' physics reasoning abilities.

    By Ansari; Ali; Sun; Haoran; Liu; Andy Zeyi; Jabbour; Mark; Ding; Yongshan; Girvin; Steven; Yu; Ismail-Beigi; Sohrab; Kubica; Aleksander; Miller; Owen D; O'Hern; Corey; Ozolins; Vidvuds; Poland; David; Stone; A Douglas; Bosch; Frank C van den; Wright; Logan; Akbari; Navid; Antu; Santanu; Cai; Kangle; Calabrese-Day; Andrew; Wuttig; Mateo Cárdenes; Cheng; Meng; Chiang; Barry T; Ghorashi; Gu; Shouzhen; Huang; Haoyang; Zhibo; Kienesberger; Lukas; Hantian; Lomba; Charles; Zhongling; Wenchao; McIntosh; Rohin E; McKinney; Evan; Rojkov; Ivan; Xulei; Tokayer; Yarone Meir; Umasankar; Naveen Balaji; Varma; Mira; Wang; Leda; Qimin; Tyler; Wei; Haoyu; Yang; Jinming; Zhao; Jinchen; Sherlock Tingrui; Zheng; Qinyuan; Zou; Jay S; Baker; Lucas; Cohan; Arman; Sous; John
  12. 012Hacker NewsSEP · 16English

    How to Build Effective Evals for AI Agents

    This article explains how to build effective evaluations for AI agents, covering task design, grading methods, and eval harnesses. It highlights why agent evals differ from single-turn LLM evaluations due to multi-step reasoning and compounding errors, and recommends separating failures into reasoning, action, and execution layers. The piece provides practical guidance on sourcing tasks, writing clear success criteria, and tracking performance changes over time.

    By Bala Priya C
  13. 013Hacker NewsSEP · 16English

    Fractal basins trap latent reasoning

    Researchers demonstrate that reasoning models in AI exhibit transient chaos and fractal basin structures that increase with task difficulty, causing longer reasoning times on harder problems. This phenomenon occurs when reasoning becomes trapped near saddle points corresponding to nearly-correct solutions, explaining why frontier models spend more computational effort on complex tasks like theorem solving and puzzle solving.

    By Lai; Jeffrey; Bao; Anthony; Quinn; John; Gilpin; William
  14. 014Hacker NewsSEP · 16English

    Qwen3.8 Max – Cost per task higher than Astra on Artificial Analysis

    Qwen3.8 Max (0902) by Alibaba achieves an above-average Intelligence Index score of 45 but generates excessive output tokens and operates at slower speeds than comparable models. Despite competitive pricing at $2.00 per 1M input tokens and $6.00 per 1M output tokens, its cost per task ($4934.79) exceeds alternatives like Astra due to verbose output generation.

    By mydreamof
  15. 015Hacker NewsSEP · 16English

    Salesforce and Nvidia launch Koa, a CRM reasoning model

    Salesforce and NVIDIA announced Koa, a CRM reasoning model built on NVIDIA Nemotron and trained with 27 years of Salesforce enterprise data to help AI agents handle complex workflows. Koa matches leading model performance with three times fewer errors on CRM tasks and runs entirely within Salesforce's infrastructure, with no customer data used in training. The companies are also bringing Nemotron models to Missionforce for government and regulated organizations requiring secure, private deployments.

    By Salesforce Newsroom
  16. 0169to5GoogleSEP · 16English

    Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail, & Keep

    Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models, designed for intuitive AI conversation. Extended Thinking offers enhanced reasoning for complex tasks across Gmail Live, Docs Live, and Keep Live, while the base model provides conversational intelligence with visual processing and support for 97 languages. Both models demonstrate strong benchmark performance in speech quality and agentic task completion.

    By Abner Li
  17. 017Hacker NewsSEP · 16English

    Atlas-Finance: Evaluating AI Agents Inside a Bank

    ATLAS-Finance is a new benchmark with 100 expert-level financial tasks across 13 realistic firm environments, testing AI agents on complex, ambiguous work requiring multi-party coordination and contextual reasoning. Frontier models including Claude Opus 5 achieved only 12.3% pass rate, with consistent failures in applying correct financial logic, maintaining required scope, and propagating calculated values—errors that would require senior auditing in actual banking practice.

    By cjbarber
  18. 018Hacker NewsSEP · 15English

    Gemini 3.8 Live and 3.8 Live Extended Thinking

    Google introduces Gemini 3.8 Live and 3.8 Live Extended Thinking, two new AI models designed for real-time voice conversations and complex reasoning tasks. The models deliver near real-time processing, multi-language support, and tool execution capabilities, with 3.8 Live Extended Thinking achieving top scores on Speech to Speech Quality benchmarks while maintaining competitive pricing.

    By Tom Ouyang; Malini Jaganathan
  19. 019Hacker NewsSEP · 14English

    The Deconstruction of Mathematical Reasoning

    An article arguing that large language models have zero mathematical capabilities, distinguishing between statistical pattern matching and genuine mathematical reasoning. The author contrasts LLMs with actual mathematical tools like proof assistants and calculators, explaining that mathematics requires rigorous logical inference rather than token prediction.

    By Ramkumar Ramachandra
  20. 020Hacker NewsSEP · 14English

    Categorizing AI inference: Initialization, Reasoning, Orchestration, Synthesis

    An article categorizing AI model inference calls into four functional roles: Initialization (processing task-independent context), Orchestration, Reasoning (resolving task-relevant uncertainty), and Synthesis. The taxonomy distinguishes inference's functional purpose beyond aggregate token metrics, enabling more precise measurement of inference value versus waste in agent systems.

    By Jeff Auriemma