source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
TUESDAY, SEPTEMBER 22, 2026
Hacker News3640X 主题热门3508CNBC729to5Mac57MacRumors56YahooFinance47Kotaku37IGN34Verge32aihot26TechCrunch25Gematsu249to5Google23NintendoLife21BusinessInsider18Eurogamer17Polygon15PushSquare14Engadget13NPR13Guardian13WarhammerCommunity13Fortune12NBC12Notebookcheck12Wccftech12AndroidAuthority11FoxBusiness11USAToday11ArsTechnica10CoinDesk10TechPowerUp10ABC9AppleInsider9SeekingAlpha9bgr8CBS8Gizmodo8PureXbox8Yahoo8CNN7MotleyFool7Mashable7NintendoEverything7PetaPixel7VideoCardz7GSMArena6Investor'sBusinessDaily6Conversation6WIRED6BleepingComputer5CNET5Fox5GameInformer5Pokemon5RockPaperShotgun5SamMobile5VideoGamesChronicle5AndroidCentral4AndroidPolice4DigitalFoundry4Motor14XBOXWire4NewYorkPost4RPGSite4SlashGear4TechSpot4Variety4WindowsCentral4WSB-TV4Aftermath3AlJazeera3BellofLostSouls3Deadline3Hodinkee3HuffPost3MP1st3PlayStationLifeStyle3SeattleTimes3Register3Tom'sGuide3TweakTown3404Media280Level2ABC7LosAngeles2BleedingCool2CTech2CanonRumors2DW2EventHubs2FratelloWatches2Futurism2GameRant2GearPatrol2HouseDigest2InsiderGaming2LosAngelesTimes2Lifehacker2Nature2PCMag2qz2SouthChinaMorningPost2Intercept2TimeExtension224/7WallSt.16abcPhiladelphia1BusinessInsiderAfrica1Alternet1AndroidHeadlines1AOL1ArizonaSports1Benzinga1BikeRadar1Billboard1BloodyDisgusting1Boston1Bungie1BusinessTimes1BuzzFeed1CalMatters1CarBuzz1cbn1ChromeUnboxed1Chron1ClaimDepot1ColoradoSun1ChristianScienceMonitor1Currently1CyberSecurityNews1Cyclingnews1DailyDownforce1DailyKos1DaringFireball1DarkHorizons1DCRainmaker1Decrypt1Defector1Defense1denver71DenverPost1DigitalCameraWorld1DirtonDirt1Draftsim1DroidLife1DSOGaming1DualShockers1empireonline1erictopol.substack1Euronews1Fangoria1FOX191DetroitFreePress1GameDeveloper1GamesIndustry.biz1GamingOnLinux1AAAGasPrices1GeekWire1GeekyGadgets1Global1Hackaday1HollywoodReporter1HoustonChronicle1Independent1Invezz1KITCO1Magic:Gathering1MakeUseOf1Maxroll1Mercury1MLive1MonochromeWatches1MorningBrew1MortgageDaily1Motorsport1MyNintendo1Newser1SemiAnalysis1Newsshooter1Newsweek1nrn1OneMileataTime1OregonPublicBroadcasting1OregonLive1PageSix1PokeBeach1politico.eu1QuantaMagazine1Road&Track1RoadtoVR1RockstarINTEL1Salon1CultureMapSanAntonio1ScienceAlert1Semafor1SFGATE1YahooSingapore1Slate1SlippedDisc1Space1SpaceNews1YahooTech1DailyBeast1DailyMeal1Drive1Hacker1Hindu1NextWeb1Times1TimesofIndia1TimesUnion1TMZ1YahooFinanceUK1UploadVR1VisualCapitalist1WindowsLatest1WKYT1WOWT1YourTango1
  1. 001Hacker NewsSEP · 22English

    AMD's random number generator can't generate a 0?

    A user reports that AMD's random number generator appears unable to produce the value 0 across multiple test runs on two AMD processors, while noting significant performance variations in RNG implementations across different CPU generations from Intel and AMD. The user awaits feedback from AMD's escalated internal investigation.

    By BruceEel
  2. 002Hacker NewsSEP · 22English

    Laya vs Jev head-to-head on identical inputs

    A head-to-head benchmark compares Laya and Jev language models on 751 identical test cases across 9 suites, finding Jev outperforms on multi-class and non-English tasks (intent 0.975 vs 0.725, toxic 1.000 vs 0.767) while Laya wins on agnews and mnli with zero cost and lower latency (180–660 ms vs 925–1068 ms). Emotion classification is weak on both models near 0.55 accuracy; gating at 0.85 confidence keeps 58% of Laya traffic at 0.878 accuracy and 78% of Jev at 0.917.

    By Instax-Dutta
  3. 003Hacker NewsSEP · 22English

    ArtifactBench: Evaluating AI Music Detectors Under Distribution Shift

    ArtifactBench is a new evaluation framework for AI-generated music detectors that accounts for distribution shifts, generator lineage, and real-world audio variations. The benchmark reveals significant performance gaps across detectors, with ArtifactNet achieving 0.982 AUROC while public detectors like Deezer's drop to 0.761 under shifted conditions.

    By Oh; Heewon
  4. 004Hacker NewsSEP · 22English

    Android Bench 2.0 – Long Horizon Android Development Benchmark

    Android Bench 2.0 is a benchmark designed to measure LLM capabilities in AI-assisted Android development, addressing gaps in existing benchmarks. Long-horizon tasks achieved a 28% pass rate with 82.2% average completion over 7.9 hours at $375.7 average cost, while per-task results showed declining performance metrics across more complex scenarios.

    By bentrengrove
  5. 005Hacker NewsSEP · 21English

    Hemmingway-1: The AI that writes like a person

    Hemmingway-1 is a 27-billion-parameter open-weight AI model optimized for everyday writing tasks like emails and messages. It outperformed larger models including GPT-6 Astra and Fable 5.1 in blind evaluations, scoring 26 points higher than competitors on human-likeness and excelling at practical writing requests like financial and administrative correspondence.

    By nateb2022
  6. 006WccftechSEP · 21English

    An M6 Pro Listing On Geekbench 7 Is Getting A Lot Of Hype For Having Higher Scores Than The 18-Core M5 Max, But Don’t Believe That These Are Legit Results

    An M6 Pro listing on Geekbench 7 showing superior performance to the M5 Max is likely fraudulent, according to Geekbench creator John Poole who identified internal inconsistencies. Apple is reportedly skipping the M6 Pro and M6 Max in favor of the M7 Pro and M7 Max, making the benchmark results suspect.

    By Omar Sohail
  7. 007WccftechSEP · 21English

    Snapdragon 8 Elite Extreme Gen 6’s Engineering Prototype Video Surfaces Before Official Launch, Showing A New Packaging, Improved Benchmark Scores & More

    A Chinese content creator leaked an engineering prototype video of Qualcomm's Snapdragon 8 Elite Extreme Gen 6, revealing its 2nm design, new CPU configuration, and Geekbench 6 scores. While the chipset shows strong multi-core performance, Apple's A20 Pro remains faster in single-core benchmarks. The prototype features a new Heat Pass Block for thermal management ahead of the Snapdragon Summit.

    By Omar Sohail
  8. 008Hacker NewsSEP · 21English

    I'm afraid of spiders. So I made AI look at 2k of them

    The author tested nine AI models on 2,000 spider photos to evaluate their species identification accuracy. Gemini 3.8 Flash achieved the highest score at 49.85%, while most models correctly identified the broader spider family in 87–92% of cases even when missing the exact species.

    By Dawid Kiełbasa
  9. 009Hacker NewsSEP · 21English

    Exploring the scalable matrix extension of the Apple M4 processor

    The Apple M4 chip in the 2024 iPad Pro is the first public device supporting ARM's scalable matrix extension (SME), enabling direct low-level programming of matrix hardware for improved performance in scientific and machine learning tasks. A researcher has created microbenchmarks to explore M4 SME capabilities, measuring compute throughput and memory transfer rates across various matrix and vector operations.

    By Tzakharko
  10. 010Hacker NewsSEP · 21English

    Masked LFW (MLFW) Database

    Masked LFW (MLFW) is a face recognition database created by adding realistic masks to images from the Cross-Age LFW database to evaluate how masks impact facial recognition systems. State-of-the-art models show 5-16% accuracy decline on MLFW compared to unmasked images, addressing performance gaps in real-world masked face verification scenarios.

    By teleforce
  11. 011Hacker NewsSEP · 20English

    A design brief as one image: can a 27B vision model build the animation?

    A 27B vision model was tasked with building canvas animations from design briefs provided as single PNG images. Six quantized versions of Qwen3.8-27B and two Bonsai variants were evaluated across three design cards using 17 objective checks per page, with results showing Opti performing competitively with Q4_K_M despite 28% smaller file size, while Bonsai struggled with JavaScript errors and token budget exhaustion.

    By airylizard
  12. 012Hacker NewsSEP · 20English

    Jev Jailbreak Benchmark

    TypeSafe's Jev jailbreak detection model is benchmarked against four shipped detectors including Meta's on 7,803 labeled messages and 296 conversations. Jev wins on the curated benchmark and newest attack set but loses on two older benchmarks, with probability scores not matching documentation claims. The evaluation measures single-message and pre-scripted multi-turn attacks, missing adaptive defenses and multi-turn campaign attacks that evade detection by design.

    By Mike Ramos
  13. 013Hacker NewsSEP · 20English

    Show HN: Will Jev pull the lever in the trolley problem?

    A Hacker News post introduces a benchmark tool that tests how the AI model Jev responds to trolley problem scenarios, where users place items on railroad tracks and Jev decides whether to pull a lever. The creator reports that Jev performs well at the task and notes it hasn't pulled the lever away from villains or geese in their testing.

    By lanyard-textile
  14. 014Hacker NewsSEP · 20English

    Seven, Forever

    A project builds a benchmark to detect gaps between software documentation and actual code behavior using automated tools. It provides baselines for verifying whether code matches its claimed functionality and invites contributions of examples, particularly subtle cases and security-related discrepancies.

    By o2zer0cool
  15. 015Hacker NewsSEP · 19English

    Brood War Bench

    AI models were benchmarked playing StarCraft: Brood War, with Codex Astra performing best but all models remaining at beginner level. Codex excelled at disruptive tactics like harassing workers but struggled with sustained production, while Grok spent excessive time reasoning without taking sufficient actions. Fable showed the most genuine game understanding by attempting economy and tech progression.

    By Ben Swerdlow
  16. 016Hacker NewsSEP · 19English

    M6 Pro Chip Result on Geekbench Is Likely Fake

    A Geekbench result purporting to show an unreleased M6 Pro chip's performance likely contains fabricated data, according to Geekbench creator John Poole, who identified internal inconsistencies. This aligns with previous reports that Apple plans to skip M6 Pro and M6 Max chips.

    By Joe Rossignol
  17. 017Hacker NewsSEP · 19English

    Tin: full-text search for Postgres

    Tin is a new full-text search extension for Postgres that supports Boolean expressions, fuzzy matching, phrase queries, and BM25 scoring while handling joins, updates, and replication. The developers benchmarked Tin against existing alternatives using large corpora including Wikipedia and Stack Exchange data, finding it significantly faster across conjunction, disjunction, and phrase query workloads.

    By ksec
  18. 018Hacker NewsSEP · 19English

    We found defects in 37 of DeepSWE's 113 tasks

    Researchers identified defects or ambiguous requirements in 37 of DeepSWE's 113 tasks (32.7%), a benchmark used to evaluate AI models like GPT-6 Astra and Fable 5. Issues included hidden tests causing build failures, assertions rejecting valid output, and unspecified requirements. Fixing confirmed defects raised measured pass rates by 4–6 percentage points, raising questions about benchmark reliability.

    By lebek
  19. 019Hacker NewsSEP · 19English

    Stepfun Step 5 Preview (LLM): On AA Pareto frontier

    StepFun's Step 5 Preview is a proprietary reasoning model with 600B parameters released September 18, 2026, scoring 44 on the Artificial Analysis Intelligence Index with competitive pricing of $1.00 per 1M input tokens and $2.70 per 1M output tokens. The multimodal model supports text and image inputs, offers a 1M token context window, and ranks well above average in intelligence compared to similarly priced models.

    By AnodicElegy
  20. 020TechPowerUpSEP · 19English

    Control Resonant Performance Benchmark Review - 35 GPUs Tested

    A performance benchmark review for Control Resonant tested across 35 GPUs. The page indicates an automated bot check is in progress.