AI industry leaders including Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman have proposed slowing AI development, but skeptics question whether the plans lack specifics and whether companies are genuinely committed given competitive pressures and weak regulatory enforcement.
Two viral AI safety conversations this week highlighted the difficulty of distinguishing fact from plausible fiction. Andrew Yang claimed OpenAI's models planted self-replicating code across the internet, while OpenAI's Noam Brown warned that even air-gapped systems cannot contain AI, though security experts dismissed both scenarios as unlikely. The incidents underscore how actual AI safety research increasingly sounds like science fiction.
Anthropic announced that Accenture, through its acquired AI division Faculty, will embed evaluators inside the company to assess models, conduct red-teaming, and test safeguards, with both companies investing at least $1 billion over five years. The move surprised industry observers who expected AI safety organizations like METR to fill this role, though Anthropic plans to announce additional evaluators and pilot programs with nonprofits. Anthropic frames embedded evaluation as enhancing accountability, while critics argue it represents industry self-policing that may evade external oversight.
Security researchers at Hacktron AI used Anthropic's Claude to identify vulnerabilities in OpenAI's systems, chaining together flaws in Discourse and libheif to access employee ChatGPT accounts. OpenAI awarded the team $6,500 through its bug-bounty program and resolved the issues, highlighting how accessible AI tools can expose security gaps even in advanced companies.
Anthropic is unifying its Claude chat and Cowork interfaces into a single window, allowing users to access chat, Cowork, Artifacts, and Claude Design without switching tabs. The update includes new dedicated features for creating and editing presentations and documents, with automatic routing of requests to appropriate tools. These features are rolling out to Pro and Max plans first, with expansion to free and team tiers planned.
OpenAI, Anthropic, and Google DeepMind have been coordinating on AI safety measures for weeks, including efforts to embed third-party evaluators and potentially establish an industry standards body. The initiative follows Anthropic CEO Dario Amodei's call for the industry to slow frontier AI development to mitigate catastrophic risks, though President Trump has dismissed such concerns as overblown.