[also see AI202x page for resources, related presentations,
fun video and my SciFi take on this at JenAI
John Daniels talk on AI Stewardship (video link) is also relevant]
In May, 2026 OpenAI started testing a new AI Model, with thousands of agents, each in its own sandbox, assigned numerous tasks including 898 "capture the flag" puzzles from "ExploitGym" testing aspects of cybersecurity. 198 of these had not been solved by any prior agents, and might not be possible.
By July 4th, OpenAI found the agents had broken out of their sandboxes, developed a communications system, started operating as a "collective" or "swarm" and overloaded one of the OpenAI systems.
By July 6th OpenAI "fixed" the zero day flaws exposed by the May test and restarted the tests with updated agents and lower guardrails.
By July 16th Hugging Face detected and announced that their systems had been compromised by an agent swarm
By Oct 2026, it is evident that many AI Agent systems have managed to reach the Internet, some hacking into 3rd party systems, starting in May 2026.
Hugging Face
Aug 1 - released a longer technical write up of the incident: https://huggingface.co/blog/agent-intrusion-technical-timeline
Open AI
BlackHat Presentation https://www.youtube.com/watch?v=87DyyMV0kCY&t=997s
July 21, 2026 OpenAI and Hugging Face partner to address security incident during model evaluation https://openai.com/index/hugging-face-model-evaluation-security-incident/
METR
Aug 26, Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident https://metr.org/hugging-face-incident-report-aug-2026.pdf