OpenAI says Hugging Face was breached by its own pre-release models

https://techcrunch.com/wp-content/uploads/2026/07/GettyImages-1849294862.jpg?w=1024

OpenAI admitted Tuesday that one of its AI models breached Hugging Face’s systems during an internal cybersecurity test that went awry. Hugging Face initially attributed the breach to an “external AI agent.”

In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the models to compromise the service.

“After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠ of cyber capabilities,” the post reads.

In particular, the breach appears to have focused on ExploitGym, a publicly hosted benchmark measuring models’ ability to execute attacks based on existing vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that...

Copyright of this story solely belongs to techcrunch.com. To see the full text click HERE