OpenAI documents a real AI-agent attack kill chain
28 Aug 2026 · AI Security · CISO Careers
OpenAI / METR 1,300+ agent transcripts · 70,000+ messages · SSRF, RCE, lateral movement
OpenAI, with METR and Redwood Research, has published a detailed investigation into the Hugging Face incident, tracing how a group of autonomous agents worked their way across a target environment.
What the reports document
The agents coordinated over a shared channel, reached network access they were not meant to have, discovered credentials, exploited a vulnerability, achieved remote code execution, and moved laterally across the environment. The agents did not "break into the internet." It was a deliberately permissive research environment. Humans have been doing this for decades. Just slightly slower.
It is an unusually detailed look at an actual AI-agent kill chain, rather than another theoretical AI-threat scenario.
How the agents coordinated
Read as a whole, the reports are less an exploit write-up than a transcript. The agents planned, split the work, and corrected one another as they went. The interesting part is what the kill chain looked like. Give several agents a shared channel and you get something close to a developer Slack, except they write working code on the first try.
Where the agents failed
The same transcripts show a great deal of failure. Plenty of tasks went nowhere. For instance, none of them produced the assessment report, with slides, to present to the executives.
For defenders
A real kill chain to study beats a hypothetical one every time, and this one comes with the transcripts attached. The most uncomfortable finding may be the simplest one: the investigation took exponentially more time than the attack itself.
See everything. Trust nothing.Keep your Eyes Wide Shut, CISO.
Sources · OpenAI, The Hugging Face incident and the road ahead ↗ · METR, OpenAI Hugging Face incident investigation ↗ · Redwood Research
← All dispatches