source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
SUNDAY, OCTOBER 11, 2026
Hacker News1693X 主题热门1522ChainCatcher111CNBC75PANews67Coinness53Channel News Asia49SouthChinaMorningPost35Crypto.news30YahooFinance30Verge22CoinDesk21Kotaku20Cointelegraph19aihot17Variety17IGN169to5Mac15MacRumors15Decrypt12TechCrunch129to5Google11AMBCrypto11DW11TechPowerUp11Guardian11Eurogamer10FoxBusiness9NintendoLife9Wccftech9BBC World8BusinessInsider7VideoCardz7BitcoinMagazine6Engadget6NBC6Polygon6WarhammerCommunity6ArsTechnica5CNET5Gematsu5Gizmodo5Investor'sBusinessDaily5XBOXWire5Register5bgr4CBS4CNN4Futurism4GSMArena4Mashable4Notebookcheck4NYT4PushSquare4USAToday4AlJazeera3AppleInsider3MotleyFool3Fortune3Fox3PokémonGOHub3SeekingAlpha3Hacker3VideoGamesChronicle3BleepingComputer2BostonGlobe2DigitalFoundry2DroidLife2DSOGaming2Euronews2KSL2Lifehacker2MyNintendo2Nature2Newser2NPR2SeattleTimes2Yahoo2WindowsCentral2WIRED2Yahoo224/7WallSt.16abcPhiladelphia1ABC7NewYork1ABC1AlineaInsightnewsletter1AndroidCentral1AndroidPolice1AOL1NikkeiAsia1Bank of England1Barron's1Beebom1BloodyDisgusting1CFTC1ChromeUnboxed1Cleveland1CreativeBloq1EventHubs1Federal Reserve1DetroitFreePress1FTC1GameDeveloper1GameFile1GameRant1GeekWire1HotHardware1HouseDigest1Independent1InsiderGaming1InvenGlobal1KTLO1KUTV1MP1st1NBC5Chicago1NBCSports1MicrosoftSource1NintendoWire1CrudeOilPricesToday1OMG!Ubuntu1PaulKrugman1PCGamer1PennLive1Phoronix1Pocket-lint1Pokemon1PureXbox1OutlookRespawn1RetractionWatch1Road&Track1RockPaperShotgun1RPGSite1ScienceAlert1SEC1SFGATE1YahooFinanceSingapore1SpaceNews1Syracuse1TimesSquareChronicles1YahooTech1Hill1Outerhaven1Times1Tom'sGuide1TopGear1TweakTown1YahooUK1OutsideMagazine1VGChartz1EdZitron'sWhere'sYourEdAt1WolfStreet1YankoDesign1ZDNET1
  1. 001Hacker NewsOCT · 08English

    Sudo L7 – A benchmark that measures the judgment behind good engineering

    Anthropic introduced sudo L7, a benchmark measuring staff-level software engineering judgment beyond basic coding ability. It evaluates 60 tasks from real production codebases across security, reliability, and architecture, grading agents on functional correctness, architectural decisions, risk awareness, and communication—not just whether tests pass. Claude Opus 5.5 and Sonnet 5.5 completed roughly 45% of tasks, revealing larger differences in engineering behavior beyond overall scores.

    By No items found