source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
WEDNESDAY, SEPTEMBER 23, 2026
Hacker News3805X 主题热门3681CNBC75YahooFinance579to5Mac53MacRumors51aihot50IGN39Kotaku39Verge39TechCrunch28Gematsu22NintendoLife22AndroidAuthority219to5Google20Eurogamer20BusinessInsider19Engadget16WarhammerCommunity15PushSquare14Guardian14USAToday14Polygon13Wccftech13ArsTechnica12AppleInsider11Fortune11Gizmodo11Investor'sBusinessDaily11NBC11NPR11FoxBusiness10Notebookcheck10SeekingAlpha10TechPowerUp10CBS9CoinDesk9ABC8bgr8CNN8MotleyFool8Mashable8NintendoEverything8VideoCardz8CNET7GSMArena7PureXbox7WIRED7Yahoo7AlJazeera6Fox6SamMobile6Conversation6VideoGamesChronicle6BleepingComputer5NewYorkPost5PetaPixel5TechSpot5Tom'sGuide5AndroidPolice4Deadline4GameInformer4XBOXWire4PlayStationLifeStyle4Pokemon4RockPaperShotgun4RPGSite4SlashGear4Variety4WindowsCentral4WSB-TV4404Media380Level3Aftermath3AndroidCentral3BellofLostSouls3DigitalFoundry3DroidLife3EventHubs3GamesIndustry.biz3HollywoodReporter3HuffPost3Motor13MP1st3Nature3PokeBeach3SouthChinaMorningPost3SeattleTimes3TimeExtension36abcPhiladelphia2ABC7LosAngeles2AZFamily2Benzinga2CanonRumors2Currently2DCRainmaker2DigitalCameraWorld2FratelloWatches2Futurism2GearPatrol2HouseDigest2InsiderGaming2LosAngelesTimes2Lifehacker2Newsweek2PCMag2qz2SFGATE2SimsCommunity2Slate2Hacker2Register2TweakTown2YahooFinanceUK2WhatHi-Fi?224/7WallSt.1BusinessInsiderAfrica1Alternet1AndroidHeadlines1AOL1ArizonaSports1BikeRadar1BloodyDisgusting1Boston1BostonGlobe1BusinessTimes1BuzzFeed1CTech1CalMatters1CarBuzz1cbn1ChromeUnboxed1Chron1ClaimDepot1ColoradoSun1Skin.ClubCommunity1ChristianScienceMonitor1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DarkHorizons1Decrypt1Defector1Defense1denver71DenverPost1Designboom1DirtonDirt1Draftsim1DSOGaming1DW1empireonline1GameGPU1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1GameDeveloper1GameRant1GameWorldObserver1GAMINGbible1GamingOnLinux1AAAGasPrices1GeekyGadgets1Global1Gothamist1Hackaday1Hackster.io1Hodinkee1HoustonChronicle1Independent1InterestingEngineering1Invezz1Jalopnik1KITCO1MacObserver1Magic:Gathering1MakeUseOf1Mashed1Maxroll1Mercury1MLive1MonochromeWatches1MorningBrew1MortgageDaily1Motorsport1MyNintendo1BloombergLaw1SemiAnalysis1Newsshooter1NintendoWire1nrn1CrudeOilPricesToday1OneMileataTime1OregonPublicBroadcasting1PageSix1politico.eu1QuantaMagazine1Realtor1Road&Track1RockstarINTEL1Salon1CultureMapSanAntonio1SeattleRed1Semafor1SanFranciscoChronicle1SimpleFlying1SlippedDisc1SoraNews241Space1SpaceNews1YahooTech1the5krunner1DailyBeast1DailyMeal1Drive1Hindu1Intercept1Times1TimesofIndia1TimesUnion1TMZ1TODAY1UploadVR1VisualCapitalist1WHYY1WindowsLatest1WKYT1WOWT1YGOrganization1YourTango1
  1. 001Hacker NewsSEP · 23English

    GPT-6 Astra has gained the ability to drive a car

    GPT-6 Astra achieved 100% progress on a driving task, completing a medium-difficulty course in 5 minutes 22 seconds while staying within 4 meters of the centerline. The leaderboard compares performance across multiple AI models, with Claude Fable 5.1 reaching 45% progress and other models performing lower.

    By Aditya Ramabadran Simon Mahns Tobias Gessler Equal contribution
  2. 002Hacker NewsSEP · 23English

    Testing AIs on 68 of the hardest open Erdos problems, verified in Lean

    Researchers introduced FrontierMath Erdős (FME), a benchmark of 68 open Erdős problems formalized in the Lean proof assistant, to systematically evaluate AI capabilities in mathematics. Five AI systems were tested with a $300 budget per problem, with only GPT-6 Astra achieving 3% success and all others scoring 0%.

    By Adamczewski; Tom; Bloom; Thomas F
  3. 003Hacker NewsSEP · 23English

    Pg_chdb, fast imports from object storage to Postgres using COPY

    Pg_chdb is a PostgreSQL extension library that enables fast data imports from object storage (S3, GCS, Azure Blob) into Postgres tables using the COPY command and chDB queries. It supports multiple data formats and demonstrates consistent performance that outperforms similar extensions like pg_duckdb and pg_lake by 2-3x for CSV, JSON, and Parquet imports.

    By ClickHouse
  4. 004Hacker NewsSEP · 22English

    Show HN: Livenerf – a benchmark for whether Opus 5.5 gets nerfed

    Livenerf is a deterministic benchmark designed to detect whether Anthropic's Claude Opus 5.5 model degrades in performance after its launch on September 22, 2026. The project uses the Inspect evaluation framework to run frozen test panels and measure statistical drift over time, tracking both accuracy changes and output token counts to catch potential model degradation.

    By Ninjahawk
  5. 005Hacker NewsSEP · 22English

    Ant Group releases finance-focused Ling-3.0-flash-Fin

    Ant Group released Ling-3.0-flash-Fin, a finance-focused open weights model designed for financial research tasks like valuation analysis and report writing. The model scores 23 on the Intelligence Index and 24 on the Finance & Accounting Index, matching competitor performance while using roughly half the active parameters of comparable models.

    By gmays
  6. 006Hacker NewsSEP · 22English

    Show HN: LinearSolveBench, interesting new benchmark to discover linear solvers

    LinearSolveBench is a new benchmark designed to advance algorithms for solving large sparse linear systems and to measure AI models' ability to discover such algorithms. It includes 9,984–113,664 row matrices from FLASH magnetic-diffusion problems and spans applications across power grids, fusion energy, quantitative finance, structural mechanics, and fluid dynamics.

    By hgarud
  7. 007Hacker NewsSEP · 22English

    JevBench, a reproducible benchmark for typed decision models

    JevBench is a reproducible benchmark for typed decision models that measures Jev 1.13.0's performance across intelligence, calibration, speed, and cost metrics. The benchmark also documents classifier.dev's fast tier, which uses Jev's model with an estimated cost of $0.0033 per 1,000 decisions under a Pro plan.

    By florianstandhar
  8. 008Hacker NewsSEP · 22English

    RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

    RRSI is a method for automatically improving LLM agent systems by iteratively refining prompts, control flow, and tooling while preventing overfitting through regularization constraints. The approach uses a budget-limited proposer and a critic-pruner selector to favor reusable mechanisms, achieving significant gains on in-distribution and out-of-distribution benchmarks while reducing computational cost.

    By Xia; Peng; Han; Rujun; Wang; Zifeng; Yanfei; Zhang; Yufan; Lee; Yoonho; Huang; Chengsong; CuiZhu; Zhongying; Ming; Yifei; Yao; Huaxiu; Gokturk; Burak; Pfister; Tomas; Chen-Yu
  9. 009Hacker NewsSEP · 22English

    AMD's random number generator can't generate a 0?

    A user reports that AMD's random number generator appears unable to produce the value 0 across multiple test runs on two AMD processors, while noting significant performance variations in RNG implementations across different CPU generations from Intel and AMD. The user awaits feedback from AMD's escalated internal investigation.

    By BruceEel
  10. 010Hacker NewsSEP · 22English

    Laya vs Jev head-to-head on identical inputs

    A head-to-head benchmark compares Laya and Jev language models on 751 identical test cases across 9 suites, finding Jev outperforms on multi-class and non-English tasks (intent 0.975 vs 0.725, toxic 1.000 vs 0.767) while Laya wins on agnews and mnli with zero cost and lower latency (180–660 ms vs 925–1068 ms). Emotion classification is weak on both models near 0.55 accuracy; gating at 0.85 confidence keeps 58% of Laya traffic at 0.878 accuracy and 78% of Jev at 0.917.

    By Instax-Dutta
  11. 011Hacker NewsSEP · 22English

    ArtifactBench: Evaluating AI Music Detectors Under Distribution Shift

    ArtifactBench is a new evaluation framework for AI-generated music detectors that accounts for distribution shifts, generator lineage, and real-world audio variations. The benchmark reveals significant performance gaps across detectors, with ArtifactNet achieving 0.982 AUROC while public detectors like Deezer's drop to 0.761 under shifted conditions.

    By Oh; Heewon
  12. 012Hacker NewsSEP · 22English

    Android Bench 2.0 – Long Horizon Android Development Benchmark

    Android Bench 2.0 is a benchmark designed to measure LLM capabilities in AI-assisted Android development, addressing gaps in existing benchmarks. Long-horizon tasks achieved a 28% pass rate with 82.2% average completion over 7.9 hours at $375.7 average cost, while per-task results showed declining performance metrics across more complex scenarios.

    By bentrengrove
  13. 013Hacker NewsSEP · 21English

    Hemmingway-1: The AI that writes like a person

    Hemmingway-1 is a 27-billion-parameter open-weight AI model optimized for everyday writing tasks like emails and messages. It outperformed larger models including GPT-6 Astra and Fable 5.1 in blind evaluations, scoring 26 points higher than competitors on human-likeness and excelling at practical writing requests like financial and administrative correspondence.

    By nateb2022
  14. 014WccftechSEP · 21English

    An M6 Pro Listing On Geekbench 7 Is Getting A Lot Of Hype For Having Higher Scores Than The 18-Core M5 Max, But Don’t Believe That These Are Legit Results

    An M6 Pro listing on Geekbench 7 showing superior performance to the M5 Max is likely fraudulent, according to Geekbench creator John Poole who identified internal inconsistencies. Apple is reportedly skipping the M6 Pro and M6 Max in favor of the M7 Pro and M7 Max, making the benchmark results suspect.

    By Omar Sohail
  15. 015WccftechSEP · 21English

    Snapdragon 8 Elite Extreme Gen 6’s Engineering Prototype Video Surfaces Before Official Launch, Showing A New Packaging, Improved Benchmark Scores & More

    A Chinese content creator leaked an engineering prototype video of Qualcomm's Snapdragon 8 Elite Extreme Gen 6, revealing its 2nm design, new CPU configuration, and Geekbench 6 scores. While the chipset shows strong multi-core performance, Apple's A20 Pro remains faster in single-core benchmarks. The prototype features a new Heat Pass Block for thermal management ahead of the Snapdragon Summit.

    By Omar Sohail
  16. 016Hacker NewsSEP · 21English

    I'm afraid of spiders. So I made AI look at 2k of them

    The author tested nine AI models on 2,000 spider photos to evaluate their species identification accuracy. Gemini 3.8 Flash achieved the highest score at 49.85%, while most models correctly identified the broader spider family in 87–92% of cases even when missing the exact species.

    By Dawid Kiełbasa
  17. 017Hacker NewsSEP · 21English

    Exploring the scalable matrix extension of the Apple M4 processor

    The Apple M4 chip in the 2024 iPad Pro is the first public device supporting ARM's scalable matrix extension (SME), enabling direct low-level programming of matrix hardware for improved performance in scientific and machine learning tasks. A researcher has created microbenchmarks to explore M4 SME capabilities, measuring compute throughput and memory transfer rates across various matrix and vector operations.

    By Tzakharko
  18. 018Hacker NewsSEP · 21English

    Masked LFW (MLFW) Database

    Masked LFW (MLFW) is a face recognition database created by adding realistic masks to images from the Cross-Age LFW database to evaluate how masks impact facial recognition systems. State-of-the-art models show 5-16% accuracy decline on MLFW compared to unmasked images, addressing performance gaps in real-world masked face verification scenarios.

    By teleforce
  19. 019Hacker NewsSEP · 20English

    A design brief as one image: can a 27B vision model build the animation?

    A 27B vision model was tasked with building canvas animations from design briefs provided as single PNG images. Six quantized versions of Qwen3.8-27B and two Bonsai variants were evaluated across three design cards using 17 objective checks per page, with results showing Opti performing competitively with Q4_K_M despite 28% smaller file size, while Bonsai struggled with JavaScript errors and token budget exhaustion.

    By airylizard
  20. 020Hacker NewsSEP · 20English

    Jev Jailbreak Benchmark

    TypeSafe's Jev jailbreak detection model is benchmarked against four shipped detectors including Meta's on 7,803 labeled messages and 296 conversations. Jev wins on the curated benchmark and newest attack set but loses on two older benchmarks, with probability scores not matching documentation claims. The evaluation measures single-message and pre-scripted multi-turn attacks, missing adaptive defenses and multi-turn campaign attacks that evade detection by design.

    By Mike Ramos