source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
THURSDAY, SEPTEMBER 17, 2026
Hacker News3938X 主题热门3797CNBC84MacRumors71YahooFinance619to5Mac58Kotaku45Verge41IGN349to5Google33aihot33Gematsu31NintendoLife31BusinessInsider27Engadget26Eurogamer26TechCrunch26Guardian18Polygon17NBC16CNET14NPR14Fortune13FoxBusiness13USAToday13Wccftech13bgr12Gizmodo12Mashable12PushSquare12SeekingAlpha12Notebookcheck10WIRED10TechPowerUp9ABC8AppleInsider8CNN8Fox8GameInformer8Investor'sBusinessDaily8NewYorkPost8VideoGamesChronicle8WindowsCentral8CBS7ArsTechnica6BleepingComputer6XBOXWire6NintendoEverything6CrudeOilPricesToday6Variety6AndroidPolice5GamesIndustry.biz5PetaPixel5PureXbox5SamMobile5AlJazeera4AndroidAuthority4CoinDesk4DigitalFoundry4GameRant4GSMArena4SlashGear4Conversation4Register4WarhammerCommunity4Yahoo4CTech3ChromeUnboxed3Deadline3DW3Jalopnik3Lifehacker3Motor13Blizzard3PCMag3PCWorld3Pokemon3RockPaperShotgun3RPGSite3SouthChinaMorningPost3SeattleTimes3Space3Hacker3TweakTown3VideoCardz3WindowsLatest3ZDNET3404Media280Level2Aftermath2AndroidCentral2AOL2AwfulAnnouncing2BleedingCool2BuzzFeed2CanonRumors2CyberSecurityNews2DroidLife2DualShockers2Euronews2EventHubs2MotleyFool2FratelloWatches2Futurism2GameDeveloper2GearPatrol2Hodinkee2KITCO2LosAngelesTimes2MassivelyOverpowered2Maxroll2MP1st2MyNintendo2Nature2Newser2PaulKrugman2PokémonGOHub2RoadtoVR2SFGATE2Intercept2NextWeb2Tom'sGuide2UploadVR2YourTango2ABC111AboveLaw1ageofempires1AVClub1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Borderlands1Boston1Bungie1Yahoo!FinanceCanada1CineD1CnEVPost1comicbook1CreativeBloq1Cyclingnews1DailyDownforce1DailyKos1Defector1DenverPost1derekthompson1DigitalCameraWorld1Draftsim1CNN1flatpanelshd1FrequentMiler1GAMINGbible1garymarcus.substack1GeekWire1GeekyGadgets1Hackaday1HollywoodReporter1Independent1InsiderGaming1InterconnectsAI1InterestingEngineering1KSL1Lloyd'sList1WPLGLocal101Macworld1Magic:Gathering1Mediaite1Mercury1MonochromeWatches1MortgageDaily1MPR1SemiAnalysis1Newsweek1nylon.com.sg1NYT1OregonLive1PageSix1PCGamesN1politico.eu1PittsburghPost-Gazette1QuantaMagazine1qz1SammyGuru1CultureMapSanAntonio1ScienceAlert1ScientificAmerican1Semafor1YahooFinanceSingapore1YahooSingapore1SportsIllustrated1SimpleFlying1supercarblondie1YahooTech1Tedium1TelecomTalk1GameBusiness1TheGamer1TimeExtension1LongmontTimes-Call1TmoNews1TopGear1TwistedVoxel1YahooFinanceUK1UnHerd1VisualCapitalist1WhatHi-Fi?1WOWT1WPBF1WRAL1WSB-TV1YGOrganization1
  1. 001Hacker NewsSEP · 17English

    Show HN: AutoBot – live voice control for long-running AI work

    AutoBot is an agentic harness for long-running AI work that achieved 18.5% higher task completion than OpenAI's baseline and ranked #1 on AssistantBench, featuring self-improving capabilities, hierarchical memory, persistent task graphs, and local computation. It integrates with native ChatGPT on Mac and emphasizes privacy, independent validation, and durable task follow-through across complex multi-application workflows.

    By Demeyer1
  2. 002Hacker NewsSEP · 17English

    Aegis: Zero-GC 64-byte cache-aligned memory arena in C++20 (1B ops in 0.649s)

    Aegis is a zero-garbage-collection, cache-aligned memory arena in C++20 designed for high-concurrency LLM inference runtimes. It achieves 1 billion operations in 0.649 seconds with minimal heap overhead by eliminating allocator churn and lock contention, fitting token verification descriptors into 64-byte cache lines.

    By Markbgilbert
  3. 003Hacker NewsSEP · 17English

    Show HN: Proxy-benchmark – is it the proxy, the browser, or your machine?

    Proxy-benchmark is a tool that isolates which component—proxy, browser, host machine, or target—is causing request failures by running controlled tests across different engines, proxy paths, and browser configurations. The project, maintained by NodeMaven, publishes reproducible results with full parameter sets so users can diagnose network issues systematically rather than through trial-and-error.

    By Nodemaven
  4. 004Hacker NewsSEP · 17English

    What Fits (Into Few Tokens) Doesn't Overfit

    A study of LLM-driven research agents demonstrates that successful machine learning strategies remain compressible and generalizable even when reusing benchmark data. Using output and input compression tests across multiple domains, researchers find that short prompts and minimal feedback suffice to reproduce high-performance models, supporting a description-length explanation for why benchmark-driven ML avoids overfitting in practice.

    By Bertran; Martin Andres; Roth; Aaron; Wu; Zhiwei Steven
  5. 005Hacker NewsSEP · 17English

    Artificial Analysis Capability Indices v1.1

    Artificial Analysis released Capability Indices v1.1, updating domain-specific AI model evaluations across finance, legal, healthcare, engineering, and other sectors. The update incorporates stronger evaluations from Intelligence Index v4.3, adds agentic tool use benchmarks, and removes customer interaction metrics across most domains.

    By wertyk
  6. 006Hacker NewsSEP · 17English

    Show HN: Compute:Arena – Community submitted local AI benchmarks

    Compute:Arena is a community-driven platform for benchmarking local AI models across different hardware and software configurations. Users submit performance metrics for various models including Qwen, Llama, and Gemma variants running on Apple Silicon and AMD GPUs, with measurements of throughput and prompt processing speed.

    By prabod
  7. 007Hacker NewsSEP · 16English

    Dwarf Star Support for Qwen3.8 Flash Next

    A technical discussion about optimizing Qwen3.8 Flash with Dwarf Star support, comparing performance benchmarks between ds4 and llama.cpp implementations. The ds4 stack achieved 16% faster performance on identical reliability metrics, though installation complexity and stack maintenance remain concerns for deployment.

    By Antirez
  8. 008Hacker NewsSEP · 16English

    Artificial Analysis: What Is the Intelligence Index Measuring?

    Artificial Analysis's Intelligence Index, a widely-used AI model leaderboard, compresses diverse evaluation choices into a single intelligence score that obscures methodological trade-offs. The index weights agentic workloads heavily (34%), relies on saturated benchmarks like GPQA that cannot distinguish frontier models, lacks genuine coding benchmarks despite a 24% coding category, and features half its components using similar agent-execution patterns that may over-represent certain capabilities.

    By baddash
  9. 009aihotSEP · 16Chinese

    GPT-5.6 Luna 对比 GPT-6 Astra:$1.20 的模型做代码评审够用吗

    A comparison of GPT-5.6 Luna and GPT-6 Astra for code review on 50 public pull requests found Luna identified 69 verified bugs versus Astra's 92, with costs of $0.20 versus $5.66 and accuracy rates of 74% versus 96% respectively.

  10. 010Hacker NewsSEP · 16English

    Show HN: TurboBench, the Compression Lie Detector, 100 Codecs, Daily Update

    TurboBench is a compression benchmarking tool that tests 100+ codecs against the Silesia Corpus on multiple platforms (linux-aarch64 and linux-riscv64). Results show compression ratios, compression/decompression speeds, and rankings, with zstd and brotli performing well on compression ratio and misa77 excelling in decompression speed.

    By Powturbo
  11. 011Hacker NewsSEP · 16English

    Coding Agents Have Converged: Why the SWE-Bench Leaderboard Can No Longer Order

    A study audits the SWE-bench leaderboard for coding agents, finding that top entries have converged with highly overlapping success sets, making small score differences unreliable for ranking. Statistical tests show most adjacent top-thirty pairs cannot be significantly distinguished, suggesting leaderboard positions don't establish clear ordering. The authors propose reporting model-scaffold-specific results and resolution metrics instead of interpreting marginal aggregate gaps as meaningful rank differences.

    By Liu; Fengshuo; Ying; Sun; Ruize; Luo; Lie; Guo; Siyuan
  12. 012Hacker NewsSEP · 16English

    GPT-6 Astra vs. GPT-5.6 Sol: Is a 1.6x Higher Cost Worth It per Verified Bug?

    A comparison of GPT-6 Astra and GPT-5.6 Sol code review models found that despite Astra's 2.5x higher cost and superior precision (95% vs 85%), Sol identified more confirmed bugs across 50 pull requests (107 vs 91) at lower cost per bug ($0.039 vs $0.062). The study highlights that meaningful code review evaluation requires measuring both bug detection and false-positive rates, not just raw findings.

    By Aditya Jha
  13. 013Hacker NewsSEP · 15English

    Open-Source Skill Makes Claude Code 38% Cheaper and 38% Faster on Small Projects

    Product Traceability 2.0, an open-source skill for Claude Code, achieved 38% cost reduction and 38% faster build times on small projects by moving product history maintenance out of the coding agent's loop. Version 1.0 failed because it required the agent to maintain four Markdown files synchronously, consuming excessive compute; Version 2.0 separates coding work from record-keeping to preserve efficiency.

    By Vlad Mysla
  14. 014WccftechSEP · 15English

    M5 Ultra Demolishes The 96-Core Ryzen Threadripper PRO 9995WX CPU In New Geekbench 7 Scores; Obtains 11% Higher Multi-Core Score With Less Than Half The Number Of Cores

    Apple's M5 Ultra processor outperformed AMD's 96-core Ryzen Threadripper PRO 9995WX in Geekbench 7 benchmarks, achieving 11% higher multi-core scores and 46% faster single-core performance despite having only 36 cores. The M5 Ultra excels in single-threaded and memory-intensive tasks, while the Threadripper PRO dominates in highly parallelized tests like ray tracing and compilation.

    By Omar Sohail
  15. 015Hacker NewsSEP · 15English

    Breaking the Token Ceiling

    This research introduces two methods to convert token logits to byte logits and conducts a large-scale study comparing byte-based and token-based language models across scaling dimensions. The study finds that while token models perform better initially, byte models eventually surpass them with increased compute, achieving superior data efficiency and performance on downstream tasks.

    By Marathe; Kalyani; Pagnoni; Artidoro; Limisiewicz; Tomasz; Margaret; Lewis; Mike; Zettlemoyer; Luke; Iyer; Srinivasan
  16. 016Hacker NewsSEP · 15English

    Which is the better data analyst? Benchmarking ChatGPT vs. Claude

    A detailed benchmark compared Claude and ChatGPT's performance on data analysis tasks using live Zendesk support ticket data. Both models were tested identically across seven stages from discovery to self-audit, with results evaluated by a third Claude instance for factual accuracy against actual returned metrics.

    By Paul Joyce
  17. 017Hacker NewsSEP · 15English

    Mouse

    Mouse is an open source harness for long-running coding agents built on OpenCode. It passed 25 of 30 tasks on FrontierHarness Eval using Kimi K3, enforcing completion loops with verification rules to improve task completion accuracy.

    By Mousedev
  18. 018Hacker NewsSEP · 14English

    Principles for Fast Tokio Applications

    Russell shares principles for optimizing Tokio async applications based on discussions at RustConf, emphasizing the balance between fairness and batching. Key insights include determining if a performance problem actually exists before optimizing, using scheduling latency metrics for diagnosis, and yielding frequently to reduce latency while batching to improve throughput. The post notes that most performance issues stem from application code rather than Tokio itself.

    By carllerche
  19. 019Hacker NewsSEP · 14English

    Bad evals, my own: five exercises from two LLM judges

    An engineer evaluates their two custom LLM judges (brief and scout) used to filter AI news and Reddit threads, applying Dan Luu's critical evaluation method. They present five exercises highlighting inconsistencies and methodological issues discovered in their evaluation suites, including variable results across repeated runs, unclear evaluation criteria, and model-dependent performance changes.

    By Alessandro Prandini
  20. 020Hacker NewsSEP · 14English

    Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

    Researchers propose intelligence per watt (IPW) as a metric to measure how efficiently local AI models can answer real-world queries on power-constrained devices. Evaluating 20+ local language models across 1M queries, they find local models successfully answer 88.7% of queries with IPW improving 5.3x from 2023-2025, demonstrating that local inference can redistribute significant demand from centralized cloud infrastructure.

    By Saad-Falcon; Jon; Narayan; Avanika; Akengin; Hakki Orhun; Griffin; J Wes; Shandilya; Herumb; Lafuente; Adrian Gamarra; Goel; Medhya; Joseph; Rebecca; Natarajan; Shlok; Guha; Etash Kumar; Zhu; Shang; Athiwaratkun; Ben; Hennessy; John; Mirhoseini; Azalia; Ré; Christopher