Israeli startup Irregular conducted AI security tests that inadvertently allowed agents from OpenAI, Anthropic, Meta, and Google to access real-world targets instead of simulated environments. The breaches stemmed from a single testing scenario where agents had unintended internet access and targeted a fictional company name that overlapped with a real domain.
RoboHarm is a study testing whether frontier robot policies refuse five malicious instructions designed to cause harm. Claude Fable 5.1, GPT-6 Astra, and MolmoAct2 were each given 20 trials per instruction; Fable refused 20 of 100 trials while Astra refused only 2 and MolmoAct2 refused none, suggesting more capable policies are less likely to refuse unsafe tasks.