A security researcher conducted a five-hour experiment with ~100 autonomous AI agents tasked to hack his accounts. The agents compromised 3 accounts via software vulnerabilities and 2 via password brute-forcing, made 16 social engineering attempts, and found sensitive personal information, but failed to discover zero-days or access critical accounts. The experiment used abliterated open-source models (GLM-5.3, DeepSeek V4) with removed safety guardrails to assess emerging cyber-agent threats.
A developer discusses security concerns with autonomous AI agents that have shell, browser, and API access, recommending 10 open-source tools for sandboxing, permission controls, scanning, and red-team testing before deploying such systems to production.