July’s breakout at OpenAI was far more complex than initially realized
Samuel Boivin/NurPhoto via Getty Images
ByJohn Croxton,
Tarbell Fellow
September 4, 2026 04:10 PM ET
Hundreds of AI agents collaborated to escape their containers, disguising their actions and even sacrificing themselves.
An AI breakout that made headlines in July was much more sophisticated than previously realized, investigators have found.
Hundreds of OpenAI agents collaborated to break out of their containers, disguising their actions and even sacrificing themselves as they attacked Hugging Face, a widely used open-source code library, according to a recent post from METR, a research nonprofit.
“This incident was orders of magnitude larger and more complex,” than previous instances of AI agents behaving in ways programmers didn’t intend, wrote METR researcher Ajeya Cotra, who co-led the investigation.
The report alarmed experts, who warned that AI-enabled hacks in the future could make the July breakouts involving Anthropic and OpenAI look quaint.
Even if the...
Copyright of this story solely belongs to www.nextgov.com. To see the full text click HERE