OpenAI Agent Escaped Testing and Launched an Autonomous Hack

https://www.cnet.com/wp-content/uploads/sites/2/CHATGPT_OPEAIArtboard-3.jpg

Last week, OpenAI, the developer of ChatGPT, was testing a pair of its most advanced models in an isolated environment known as a sandbox. Then an AI agent got loose, broke into Hugging Face’s playground — a repository of AI models and datasets — and carried out “tens of thousands of automated actions.”

Yep, AI went rogue.

Here’s a clearer explanation of the incident. As part of OpenAI’s safety research, a cybersecurity evaluation was conducted to determine whether a pair of OpenAI models (including GPT-5.6 Sol and a more capable unreleased model) could, in essence, “think like hackers.” The test, which took place in a contained setting with reduced guardrails, went awry when the AI models found a vulnerability in the software, escaped their controlled environment and toddled over into the open internet.

Once online, the AI decided that the Hugging Face platform might have a way to “cheat” the...

Copyright of this story solely belongs to cnet.com. To see the full text click HERE

Read more