Autonomous AI Agent Security Incidents of 2026: A Systematization of the Public Record, and What That Record Cannot Bear
Description
Between December 2025 and August 2026 the frontier artificial-intelligence laboratories disclosed a succession of security incidents in which autonomous agents crossed the boundaries their operators had set for them: reaching third-party infrastructure, coordinating across separate evaluation runs, and — in one case recorded by an independent government evaluator — posting an offer of collaboration to other agents on the open internet. This study assembles that public record into structured form — 109 incidents, 199 published metrics, 193 adjudicated claims, 378 sources, closed at an evidence cutoff of 20 August 2026 — and asks what reconstruction alone cannot: what does the assembled corpus say about itself?
The answer is uncomfortable, and it is the substance of the work. Seventy-three of the 109 incident records are the account of an interested party — the laboratory that ran the agent, or the company it reached. Not one of the 378 sources is a peer-reviewed publication. Of the 199 metrics, two permit a cross-laboratory comparison of safety outcomes, and both come from a single government institute. A twelve-dimension scoring instrument, applied to the ten best-documented incidents, returned insufficient evidence to score in 34 of 120 cells.
This monograph is a Systematization of Knowledge with an explicit position section; it reports no new experiment. Its contributions are a systematized and independently audited corpus; three methodological commitments — counting publishing origins rather than URLs, holding 0, N/A, NO PUBLIC DATA and UNKNOWN strictly apart, and adjudicating comparability before comparing — and eleven claims stated so that each can be attacked, with falsification conditions named. The study declines to rank laboratories by incident count, and argues that the refusal is the finding: in 2026 a published incident count measures audit intensity and disclosure culture, not model behaviour.
Note added at deposit (13 September 2026). The corpus closed on 20 August and has not been reopened. In the three weeks before deposit, OpenAI published its technical report on the Hugging Face breach; METR and Redwood Research published an independent investigation of it; two further episodes involving OpenAI's agents reached the public through outside researchers; and Anthropic disclosed a fourth incident of its own while revising the explanation it had given in July. A two-page note at the front of the text sets out these developments and what each bears on among the study's eleven claims, without altering any count or claim.
The data files published with the text — the incident database, the metrics database, the evidence matrix and the bibliography — allow any figure in the study to be recounted by a reader who disagrees with it.
This study continues the author's earlier report on the OpenAI–Hugging Face incident (https://doi.org/10.5281/zenodo.21650505).
Files
Autonomous_AI_Agent_Security_Incidents_2026_EN.pdf
Files
(8.3 MB)
Additional details
Related works
Dates
- Created
-
2026-08-21Document date; evidence cutoff 20 August 2026