source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
FRIDAY, SEPTEMBER 18, 2026
Hacker News3730X 主题热门3563CNBC69MacRumors639to5Mac57YahooFinance53Kotaku42Verge35IGN34aihot309to5Google28NintendoLife28Gematsu27BusinessInsider25Eurogamer24TechCrunch23Engadget18Polygon16Guardian16NBC15Fortune14SeekingAlpha14USAToday14Wccftech14NPR13PushSquare13bgr12CNET12Gizmodo11Mashable11FoxBusiness10CBS9Fox9Notebookcheck9ABC8AppleInsider8GameInformer8Investor'sBusinessDaily8TechPowerUp8WIRED8ArsTechnica7PureXbox7VideoGamesChronicle7WindowsCentral7BleepingComputer6CNN6CoinDesk6XBOXWire6Variety6AndroidAuthority5GSMArena5NintendoEverything5NewYorkPost5CrudeOilPricesToday5PetaPixel5SamMobile5DigitalFoundry4GameRant4Lifehacker4Motor14Pokemon4RPGSite4SlashGear4Register4VideoCardz4Yahoo4AlJazeera3AndroidCentral3AndroidPolice3CTech3ChromeUnboxed3GamesIndustry.biz3Jalopnik3LosAngelesTimes3Blizzard3RockPaperShotgun3SouthChinaMorningPost3SeattleTimes3Space3Conversation3TweakTown3WarhammerCommunity3WindowsLatest3404Media280Level2Aftermath2AOL2AwfulAnnouncing2BleedingCool2BloodyDisgusting2BuzzFeed2CanonRumors2CyberSecurityNews2Deadline2DualShockers2DW2EventHubs2MotleyFool2FratelloWatches2GameDeveloper2GearPatrol2Hodinkee2MassivelyOverpowered2Maxroll2MP1st2MyNintendo2Nature2Newser2PCWorld2PokémonGOHub2RoadtoVR2SFGATE2Hacker2Intercept2UploadVR2YourTango2ABC111AboveLaw1BusinessInsiderAfrica1ageofempires1AVClub1Benzinga1BikeRadar1Billboard1Borderlands1Boston1Bungie1Yahoo!FinanceCanada1Chron1CineD1comicbook1CreativeBloq1Currently1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1Defector1Defense1DenverPost1DigitalCameraWorld1Draftsim1DroidLife1CNN1empireonline1Euronews1Fangoria1flatpanelshd1FOX191DetroitFreePress1FrequentMiler1Futurism1GAMINGbible1AAAGasPrices1GeekWire1GeekyGadgets1Hackaday1HollywoodReporter1Independent1InsiderGaming1InterestingEngineering1KITCO1KSL1Lloyd'sList1Macworld1Magic:Gathering1Mediaite1Mercury1MonochromeWatches1MorningBrew1MortgageDaily1Newsweek1NYT1OregonLive1PageSix1PaulKrugman1PCMag1politico.eu1PittsburghPost-Gazette1QuantaMagazine1qz1RockstarINTEL1SammyGuru1CultureMapSanAntonio1ScienceAlert1ScientificAmerican1Semafor1YahooSingapore1SportsIllustrated1SimpleFlying1Slate1supercarblondie1YahooTech1Tedium1TelecomTalk1TheGamer1NextWeb1TimeExtension1LongmontTimes-Call1TmoNews1TwistedVoxel1YahooFinanceUK1UnHerd1VisualCapitalist1WOWT1WRAL1WSB-TV1YGOrganization1ZDNET1
  1. 001Hacker NewsSEP · 18English

    Turns out, Astra cannot build manufacturing CAD yet

    Frontier AI models were tested on manufacturing CAD design tasks evaluated for geometry, editability, and manufacturability. Astra failed to achieve passing scores (60%+) across multiple task families, with some runs not completing, while performance varied significantly by task type and cost efficiency.

    By Interpret AI Inc
  2. 002Hacker NewsSEP · 17English

    Multiple providers offer free mystery model, Union Alpha

    Union Alpha, a stealth model with anonymous operators, launched free on OpenCode and OrcaRouter in September 2026. Community analysis suggests GLM-5.4 compatibility based on tokenizer fingerprinting, with benchmark performance near GPT-6 Astra levels at flash-tier pricing, though developer identity and exact specifications remain unconfirmed.

    By cdnsteve
  3. 0039to5GoogleSEP · 17English

    Android Bench 2.0 focuses on long-horizon tasks, agent evaluations

    Google released Android Bench 2.0, a benchmark for evaluating AI models on complex Android development tasks requiring multiple days to complete, such as building apps from scratch and porting cross-platform applications. The new version uses continuous scoring based on functionality, visual fidelity, and regression avoidance, with GPT-6 Astra achieving the highest pass rate at 28%. Testing revealed that models excel at writing new code and deterministic transformations but struggle with refactoring, runtime validation, and unfamiliar libraries.

    By Abner Li
  4. 004Hacker NewsSEP · 17English

    Show HN: AutoBot – live voice control for long-running AI work

    AutoBot is an agentic harness for long-running AI work that achieved 18.5% higher task completion than OpenAI's baseline and ranked #1 on AssistantBench, featuring self-improving capabilities, hierarchical memory, persistent task graphs, and local computation. It integrates with native ChatGPT on Mac and emphasizes privacy, independent validation, and durable task follow-through across complex multi-application workflows.

    By Demeyer1
  5. 005Hacker NewsSEP · 17English

    Aegis: Zero-GC 64-byte cache-aligned memory arena in C++20 (1B ops in 0.649s)

    Aegis is a zero-garbage-collection, cache-aligned memory arena in C++20 designed for high-concurrency LLM inference runtimes. It achieves 1 billion operations in 0.649 seconds with minimal heap overhead by eliminating allocator churn and lock contention, fitting token verification descriptors into 64-byte cache lines.

    By Markbgilbert
  6. 006Hacker NewsSEP · 17English

    Show HN: Proxy-benchmark – is it the proxy, the browser, or your machine?

    Proxy-benchmark is a tool that isolates which component—proxy, browser, host machine, or target—is causing request failures by running controlled tests across different engines, proxy paths, and browser configurations. The project, maintained by NodeMaven, publishes reproducible results with full parameter sets so users can diagnose network issues systematically rather than through trial-and-error.

    By Nodemaven
  7. 007Hacker NewsSEP · 17English

    What Fits (Into Few Tokens) Doesn't Overfit

    A study of LLM-driven research agents demonstrates that successful machine learning strategies remain compressible and generalizable even when reusing benchmark data. Using output and input compression tests across multiple domains, researchers find that short prompts and minimal feedback suffice to reproduce high-performance models, supporting a description-length explanation for why benchmark-driven ML avoids overfitting in practice.

    By Bertran; Martin Andres; Roth; Aaron; Wu; Zhiwei Steven
  8. 008Hacker NewsSEP · 17English

    Artificial Analysis Capability Indices v1.1

    Artificial Analysis released Capability Indices v1.1, updating domain-specific AI model evaluations across finance, legal, healthcare, engineering, and other sectors. The update incorporates stronger evaluations from Intelligence Index v4.3, adds agentic tool use benchmarks, and removes customer interaction metrics across most domains.

    By wertyk
  9. 009Hacker NewsSEP · 17English

    Show HN: Compute:Arena – Community submitted local AI benchmarks

    Compute:Arena is a community-driven platform for benchmarking local AI models across different hardware and software configurations. Users submit performance metrics for various models including Qwen, Llama, and Gemma variants running on Apple Silicon and AMD GPUs, with measurements of throughput and prompt processing speed.

    By prabod
  10. 010Hacker NewsSEP · 16English

    Dwarf Star Support for Qwen3.8 Flash Next

    A technical discussion about optimizing Qwen3.8 Flash with Dwarf Star support, comparing performance benchmarks between ds4 and llama.cpp implementations. The ds4 stack achieved 16% faster performance on identical reliability metrics, though installation complexity and stack maintenance remain concerns for deployment.

    By Antirez
  11. 011Hacker NewsSEP · 16English

    Artificial Analysis: What Is the Intelligence Index Measuring?

    Artificial Analysis's Intelligence Index, a widely-used AI model leaderboard, compresses diverse evaluation choices into a single intelligence score that obscures methodological trade-offs. The index weights agentic workloads heavily (34%), relies on saturated benchmarks like GPQA that cannot distinguish frontier models, lacks genuine coding benchmarks despite a 24% coding category, and features half its components using similar agent-execution patterns that may over-represent certain capabilities.

    By baddash
  12. 012aihotSEP · 16Chinese

    GPT-5.6 Luna 对比 GPT-6 Astra:$1.20 的模型做代码评审够用吗

    A comparison of GPT-5.6 Luna and GPT-6 Astra for code review on 50 public pull requests found Luna identified 69 verified bugs versus Astra's 92, with costs of $0.20 versus $5.66 and accuracy rates of 74% versus 96% respectively.

  13. 013Hacker NewsSEP · 16English

    Show HN: TurboBench, the Compression Lie Detector, 100 Codecs, Daily Update

    TurboBench is a compression benchmarking tool that tests 100+ codecs against the Silesia Corpus on multiple platforms (linux-aarch64 and linux-riscv64). Results show compression ratios, compression/decompression speeds, and rankings, with zstd and brotli performing well on compression ratio and misa77 excelling in decompression speed.

    By Powturbo
  14. 014Hacker NewsSEP · 16English

    Coding Agents Have Converged: Why the SWE-Bench Leaderboard Can No Longer Order

    A study audits the SWE-bench leaderboard for coding agents, finding that top entries have converged with highly overlapping success sets, making small score differences unreliable for ranking. Statistical tests show most adjacent top-thirty pairs cannot be significantly distinguished, suggesting leaderboard positions don't establish clear ordering. The authors propose reporting model-scaffold-specific results and resolution metrics instead of interpreting marginal aggregate gaps as meaningful rank differences.

    By Liu; Fengshuo; Ying; Sun; Ruize; Luo; Lie; Guo; Siyuan
  15. 015Hacker NewsSEP · 16English

    GPT-6 Astra vs. GPT-5.6 Sol: Is a 1.6x Higher Cost Worth It per Verified Bug?

    A comparison of GPT-6 Astra and GPT-5.6 Sol code review models found that despite Astra's 2.5x higher cost and superior precision (95% vs 85%), Sol identified more confirmed bugs across 50 pull requests (107 vs 91) at lower cost per bug ($0.039 vs $0.062). The study highlights that meaningful code review evaluation requires measuring both bug detection and false-positive rates, not just raw findings.

    By Aditya Jha
  16. 016Hacker NewsSEP · 15English

    Open-Source Skill Makes Claude Code 38% Cheaper and 38% Faster on Small Projects

    Product Traceability 2.0, an open-source skill for Claude Code, achieved 38% cost reduction and 38% faster build times on small projects by moving product history maintenance out of the coding agent's loop. Version 1.0 failed because it required the agent to maintain four Markdown files synchronously, consuming excessive compute; Version 2.0 separates coding work from record-keeping to preserve efficiency.

    By Vlad Mysla
  17. 017WccftechSEP · 15English

    M5 Ultra Demolishes The 96-Core Ryzen Threadripper PRO 9995WX CPU In New Geekbench 7 Scores; Obtains 11% Higher Multi-Core Score With Less Than Half The Number Of Cores

    Apple's M5 Ultra processor outperformed AMD's 96-core Ryzen Threadripper PRO 9995WX in Geekbench 7 benchmarks, achieving 11% higher multi-core scores and 46% faster single-core performance despite having only 36 cores. The M5 Ultra excels in single-threaded and memory-intensive tasks, while the Threadripper PRO dominates in highly parallelized tests like ray tracing and compilation.

    By Omar Sohail
  18. 018Hacker NewsSEP · 15English

    Breaking the Token Ceiling

    This research introduces two methods to convert token logits to byte logits and conducts a large-scale study comparing byte-based and token-based language models across scaling dimensions. The study finds that while token models perform better initially, byte models eventually surpass them with increased compute, achieving superior data efficiency and performance on downstream tasks.

    By Marathe; Kalyani; Pagnoni; Artidoro; Limisiewicz; Tomasz; Margaret; Lewis; Mike; Zettlemoyer; Luke; Iyer; Srinivasan
  19. 019Hacker NewsSEP · 15English

    Which is the better data analyst? Benchmarking ChatGPT vs. Claude

    A detailed benchmark compared Claude and ChatGPT's performance on data analysis tasks using live Zendesk support ticket data. Both models were tested identically across seven stages from discovery to self-audit, with results evaluated by a third Claude instance for factual accuracy against actual returned metrics.

    By Paul Joyce
  20. 020Hacker NewsSEP · 15English

    Mouse

    Mouse is an open source harness for long-running coding agents built on OpenCode. It passed 25 of 30 tasks on FrontierHarness Eval using Kimi K3, enforcing completion loops with verification rules to improve task completion accuracy.

    By Mousedev