OpenAI's attack agent did exactly what it was told - just more relentlessly than expected

https://www.zdnet.com/a/img/resize/a035c23a58ce8308e17a308b1984c3e9ec766e00/2026/07/23/389cae2b-9827-44ac-8c8d-c560b498fe79/hugging-face-1.jpg?auto=webp&fit=crop&height=675&width=1200

Follow ZDNET: Add us as a preferred source on Google.


ZDNET's key takeaways

  • Tests of OpenAI models led to a breach of Hugging Face systems.
  • The attack happened after OpenAI's agentic AI escaped a sandbox.
  • The threat was non-malicious, but experts expect similar incidents.

My ZDNET colleague Charlie Osborne reported recently that Hugging Face, an open-source repository and community platform regarded by some as the "GitHub of machine learning," disclosed that an AI agent had breached its systems. Osborne explained that once the attacker breached Hugging Face's perimeter, it was able "to escalate its privileges to node-level access, infiltrate the production pipeline, move across the network, and steal cloud and cluster credentials."

On Tuesday, in a post on its website, tech giant OpenAIrevealed not only that the "malicious" AI agent responsible for the breach was one of its own, but also that it viewed the attack as an...

Copyright of this story solely belongs to zdnet.com. To see the full text click HERE

Read more