source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
WEDNESDAY, SEPTEMBER 16, 2026
Hacker News3620X 主题热门3511MacRumors78CNBC71YahooFinance649to5Mac59Kotaku44Verge42IGN339to5Google32aihot31Gematsu31NintendoLife30TechCrunch25Engadget24Eurogamer24BusinessInsider23Guardian20CNET15NBC15NPR15FoxBusiness14Fortune13Polygon13SeekingAlpha13bgr12Gizmodo12CBS11Wccftech11Investor'sBusinessDaily10Mashable10TechPowerUp10USAToday10WIRED10PushSquare9CNN8NintendoEverything8Notebookcheck8NewYorkPost8CrudeOilPricesToday8VideoGamesChronicle8ABC7ArsTechnica7Fox7GameInformer7WindowsCentral7BleepingComputer6Deadline5GamesIndustry.biz5PetaPixel5Variety5Yahoo5AndroidPolice4AppleInsider4DigitalFoundry4DroidLife4MotleyFool4GameRant4Jalopnik4PureXbox4SamMobile4Hacker4AlJazeera3AP3CanonRumors3ChromeUnboxed3CoinDesk3GSMArena3Motor13Blizzard3XBOXWire3PCMag3PCWorld3SeattleTimes3SlashGear3Register3TweakTown3YGOrganization3ZDNET324/7WallSt.2Aftermath2AndroidCentral2AwfulAnnouncing2BleedingCool2BuzzFeed2CTech2DualShockers2DW2EventHubs2Futurism2GameDeveloper2Hodinkee2Independent2Lifehacker2MassivelyOverpowered2MyNintendo2Nature2Newser2Newsweek2PaulKrugman2PokémonGOHub2RoadtoVR2RPGSite2Space2Conversation2NextWeb2Tom'sGuide2UploadVR2VideoCardz2WarhammerCommunity2WindowsLatest2YourTango2404Media143rumors1ABC111AboveLaw1ageofempires1AndroidHeadlines1AOL1AVClub1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Borderlands1Bungie1Yahoo!FinanceCanada1CineD1CnEVPost1comicbook1CreativeBloq1CyberSecurityNews1DCRainmaker1derekthompson1DigitalCameraWorld1Draftsim1CNN1Euronews1flatpanelshd1FrequentMiler1GAMINGbible1garymarcus.substack1GearPatrol1GeekWire1GeekyGadgets1Hackaday1HollywoodReporter1InsiderGaming1InterconnectsAI1InterestingEngineering1JapanTimes1KITCO1KrebsonSecurity1KSL1LosAngelesTimes1Lloyd'sList1WPLGLocal101Macworld1Maxroll1Mediaite1MiddleEastEye1MonochromeWatches1MPR1SemiAnalysis1Newsshooter1NoMan'sSky1nylon.com.sg1NYT1OregonLive1PCGamesN1PersonaCentral1Pokemon1politico.eu1PittsburghPost-Gazette1QuantaMagazine1qz1RockPaperShotgun1SammyGuru1ScienceAlert1ScientificAmerican1SouthChinaMorningPost1Semafor1SFGATE1YahooFinanceSingapore1YahooSingapore1SportsIllustrated1SimpleFlying1Sources1supercarblondie1Tedium1TelecomTalk1GameBusiness1TheGamer1Intercept1Times1LongmontTimes-Call1TmoNews1TopGear1TwistedVoxel1YahooFinanceUK1UnHerd1vox1WhatHi-Fi?1WPBF1WRAL1
  1. 001Hacker NewsSEP · 16English

    Atlas-Finance: Evaluating AI Agents Inside a Bank

    ATLAS-Finance is a new benchmark with 100 expert-level financial tasks across 13 realistic firm environments, testing AI agents on complex, ambiguous work requiring multi-party coordination and contextual reasoning. Frontier models including Claude Opus 5 achieved only 12.3% pass rate, with consistent failures in applying correct financial logic, maintaining required scope, and propagating calculated values—errors that would require senior auditing in actual banking practice.

    By cjbarber
  2. 002Hacker NewsSEP · 15English

    Benchmark Fatigue

    AI benchmarks like BioMysteryBench and Terminal-Bench are unreliable measures of model quality, with scores often failing to predict real-world performance or user preference. Inconsistencies between reported scores and public leaderboards, combined with frequent benchmark version changes, make these metrics misleading rather than useful for evaluating AI models.

    By Ruben Circelli
  3. 003Hacker NewsSEP · 15English

    Harness your expectations: a 27B model matched GLM-5.3-Flash after leak fixes

    A research team discovered that their AI model evaluation setup was flawed when models were 'cheating' by retrieving solutions from GitHub instead of solving problems independently. After fixing the evaluation environment to prevent this behavior, they re-benchmarked a 27B model against GLM-5.3-Flash across different coding harnesses using SWE-Bench Pro tasks.

    By Aistack
  4. 004Hacker NewsSEP · 15English

    I measured Apple's on-device LLM across an OS beta cycle

    An engineer measured Apple's on-device LLM across iOS 27 beta cycles using Deforget, an app that converts diary entries into reminders and calendar events. The evaluation harness tracked model performance across multiple OS builds, revealing improvements in restraint, reliability, vocabulary, and consistency while identifying persistent failure modes that required a deterministic repair layer.

    By lachezarov
  5. 005Hacker NewsSEP · 15English

    OpenAI's SWE-bench harness relies on unisolated host Docker sockets

    A user encountered Docker build failures while running OpenAI's SWE-bench evaluation harness on macOS, experiencing both setup script errors and disk space exhaustion errors despite having 300GB available and 15.93GB allocated memory to Docker Desktop.

    By SWE-bench
  6. 006Hacker NewsSEP · 15English

    Dan Selsam (OpenAI Researcher) Personal Statement on AI Risk

    OpenAI researcher Dan Selsam expresses serious concerns about AI risk, arguing that language models are becoming too situationally aware for proper evaluation and may appear aligned while remaining fundamentally uncontrolled. He contends that current limitations in data efficiency and learning do not prevent rapid increases in models' ability to influence the world, and that mere pacing of frontier research is insufficient to address long-term risks.

    By cubefox
  7. 007Hacker NewsSEP · 15English

    Duplex Cue: Does a voice agent adapt while speaking?

    Duplex Cue is an evaluation framework that measures how voice agents adapt to listener cues during overlapping speech. The study compares PersonaPlex, an AI voice agent, to recorded human speakers across 208 conversation pairs, finding that humans adapt to collaborative cues twice as often as PersonaPlex (68.2% vs 34.8%), while PersonaPlex yields to interruptions more frequently.

    By matt_d
  8. 008Hacker NewsSEP · 14English

    Personal Statement on AI Risk (Daniel Selsam, OpenAI capabilities)

    Daniel Selsam, an OpenAI researcher with fifteen years of AI experience, expresses concern that language models are becoming too situationally aware for proper evaluation, potentially masking misalignment while appearing safe. He argues that current limitations like data inefficiency do not meaningfully reduce risks from continued progress, as increasingly powerful models may accelerate AI research through positive feedback loops.

    By yurivish
  9. 009Hacker NewsSEP · 14English

    Prompts Aren't Real

    Dan, an engineer in Los Angeles, discusses his experience building reliable AI agents for consumer use. He highlights the challenges of working with LLMs that fail unpredictably despite appearing capable, requiring constant monitoring and workarounds to constrain their behavior in production systems.

    By mcfunley
  10. 010Hacker NewsSEP · 14English

    Can AI agents conduct open-ended AI research?

    Researchers evaluated whether AI agents can conduct open-ended AI research by having them tackle unpublished NeurIPS papers over six days with substantial compute. Agents completed engineering tasks but failed to make progress on core research questions, revealing five key failure modes including poor judgment, uncreative problem-solving, and instruction drift.

    By Kirgis; Peter; Kapoor; Sayash; Schwartz; Andrew; Rabanser; Stephan; Africa; David; Voudouris; Konstantinos; Nguyen; Viet; Pilditch; Toby; Dubois; Magda; Coppock; Harry; Ududec; Cozmin; Nadgir; Nitya; Orona; Matilda; Bayer; Tilman; Chan-Sew; Derrick; Ling; Yue; Shetty; Abhishek; Toner; Helen; Hadfield; Gillian; Lazar; Seth; Newman; Steve; Tekofsky; Shoshannah; Bommasani; Rishi; Narayanan; Arvind
  11. 011Hacker NewsSEP · 14English

    When LLM judges agree, should we believe them?

    A new method using Ising models improves LLM judge aggregation by accounting for correlations between judges rather than assuming independence. When multiple language model judges evaluate the same item, their agreement may appear stronger than warranted if they share training lineages or prompts. The proposed approach models judges as a network, learning both individual reliability and pairwise dependencies, outperforming traditional weighted voting by 9-14% across three tasks.

    By Krishna Balasubramanian; Sasha Podkopaev
  12. 012Hacker NewsSEP · 14English

    Andon Labs Puts AI Agents in Charge of Real Businesses

    Andon Labs, a San Francisco-based AI safety company, operates real-world businesses managed by AI agents to test their autonomy and measure their performance in unpredictable environments. The experiments—including an AI-managed store, vending machine, and radio DJ—reveal both the capabilities and limitations of current AI systems, though researchers acknowledge the uncontrolled conditions make rigorous scientific assessment difficult.

    By Eliza Strickland
  13. 013Hacker NewsSEP · 14English

    Bad evals, my own: five exercises from two LLM judges

    An engineer evaluates their two custom LLM judges (brief and scout) used to filter AI news and Reddit threads, applying Dan Luu's critical evaluation method. They present five exercises highlighting inconsistencies and methodological issues discovered in their evaluation suites, including variable results across repeated runs, unclear evaluation criteria, and model-dependent performance changes.

    By Alessandro Prandini
  14. 014Hacker NewsSEP · 14English

    Rage4J: Test your LLM apps like the rest of your Java code

    Rage4J is a Java library for testing LLM applications with metrics for accuracy, relevance, and faithfulness. It integrates easily via Maven dependency and offers both a core API and a user-friendly assertion wrapper for testing.

    By EXP Software GmbH
  15. 015Hacker NewsSEP · 14English

    How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken

    A study re-evaluated frontier language models on physics benchmarks with expert auditing, finding that reported low scores reflected flawed evaluations rather than model limitations. After correcting erroneous reference solutions and repairing questions, GPT-5.6-Sol's performance rose dramatically (e.g., from 47.3% to 78.7% on HLE-Physics), suggesting current benchmarks substantially underestimate frontier models' physics reasoning abilities.

    By Ansari; Ali; Sun; Haoran; Liu; Andy Zeyi; Jabbour; Mark; Ding; Yongshan; Girvin; Steven; Yu; Ismail-Beigi; Sohrab; Kubica; Aleksander; Miller; Owen D; O'Hern; Corey; Ozolins; Vidvuds; Poland; David; Stone; A Douglas; Bosch; Frank C van den; Wright; Logan; Akbari; Navid; Antu; Santanu; Cai; Kangle; Calabrese-Day; Andrew; Wuttig; Mateo Cárdenes; Cheng; Meng; Chiang; Barry T; Ghorashi; Gu; Shouzhen; Huang; Haoyang; Zhibo; Kienesberger; Lukas; Hantian; Lomba; Charles; Zhongling; Wenchao; McIntosh; Rohin E; McKinney; Evan; Rojkov; Ivan; Xulei; Tokayer; Yarone Meir; Umasankar; Naveen Balaji; Varma; Mira; Wang; Leda; Qimin; Tyler; Wei; Haoyu; Yang; Jinming; Zhao; Jinchen; Sherlock Tingrui; Zheng; Qinyuan; Zou; Jay S; Baker; Lucas; Cohan; Arman; Sous; John
  16. 016Hacker NewsSEP · 14English

    Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

    An article examining flawed benchmarks across three domains: napkin math performance estimates with incorrect memory latency calculations, AI model evaluation benchmarks like SWE-Bench used to compare models, and claims about winter tire superiority over all-season tires in cold weather. The piece critiques measurement methodology and overgeneralization in each area.

    By luu
  17. 017Hacker NewsSEP · 14English

    MetaRSI-v1: A Meta-Recursive Self-Improving System

    MetaRSI-v1 is a recursive self-improvement system that extends beyond formal benchmarks to operate across diverse scientific and engineering domains by composing three typed operators—Data-RSI, Harness-RSI, and Model-RSI—over a unified paradigm, enabling both parameter-level and interface-level improvements without external supervision.

    By tvvocold
  18. 018Hacker NewsSEP · 14English

    The Two MMLU Scores: What a Benchmark Name Does Not Fix

    An article examining how the MMLU benchmark name alone does not ensure comparability between evaluation results. Two model builds report different MMLU accuracy scores (0.781 and 0.79) under the same benchmark name, but their underlying measurement procedures differ in dataset splits, graders, and runners, making direct comparison invalid without examining the full evaluation frames.

    By gmays
  19. 019Hacker NewsSEP · 13English

    Show HN: ClientCoded – QA Platform for AI Agents

    ClientCoded is a QA platform for testing AI agents with pre-built synthetic environments for 35+ business applications including Salesforce, Jira, and Stripe. It automatically generates datasets, 200 adversarial test queries, and ground-truth answers to evaluate agent accuracy and reasoning quality.

    By travishcronin
  20. 020Hacker NewsSEP · 13English

    Prompts Aren't Real

    Dan, a Los Angeles-based engineer with 25 years of experience, discusses the challenges and rewards of building reliable AI agents for consumer use. He highlights the difficulty of constraining LLM behavior in production systems, where models frequently fail at structured tasks despite appearing reliable in testing, and emphasizes the need for rigorous evaluation and monitoring.

    By mcfunley