AI Agents Taking Unsanctioned Action During Cyber Testing
The UK’s AI Security Institute (AISI) has disclosed a security incident in which frontier AI agents took unsanctioned actions against real people and organisations during a controlled cybersecurity evaluation.
This included attempts to socially engineer software maintainers and insert malicious code into an open-source project.
The incident happened during routine cyber capability testing between 25 and 28 July, when researchers gave AI agents access to the public internet and deliberately disabled some cyber safety mechanisms to better understand how the models behave under less restrictive conditions.
According to AISI, researchers ran a cybersecurity challenge 122 times across seven frontier models.
In 10 runs, agents took autonomous actions outside the intended scope of the evaluation, resulting in 19 unsanctioned actions directed at real people and organisations. Seventeen of those actions involved Anthropic’s Mythos 5 model, while two involved OpenAI’s GPT-5.6 Sol with cyber classifiers disabled.
The worst incident saw an agent...
Copyright of this story solely belongs to informationsecuritybuzz.com. To see the full text click HERE