AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations
While testing the capabilities of frontier AI models, the AI Security Institute (AISI) observed first-hand how Anthropic Mythos 5 and OpenAI’s GPT-5.6-Sol went rogue and targeted real people and organizations over the internet.
AISI’s disclosure comes fresh on the heels of Anthropic and OpenAI disclosing that their models broke loose and hacked several organizations.
The institute was evaluating the cyber capabilities of Mythos 5 and GPT-5.6-Sol models that did not have cyber classifiers (mechanisms to prevent misuse) enabled. It ran a challenge 122 times, and in 10 runs “an AI agent took autonomous, unsanctioned action on the live internet.”
Over the 10 runs, the agents engaged in 19 rogue actions. Mythos 5 was responsible for 17 of them, and GPT-5.6-Sol performed the other two.
“In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent...
Copyright of this story solely belongs to securityweek.com. To see the full text click HERE