AI researchers let models off the leash – then watched as they tried to add malware to a FOSS project

https://image.theregister.com/252578.jpg?imageId=252578&x=0&y=0&cropw=100&croph=100&panox=0&panoy=0&panow=100&panoh=100&width=1200&height=683

AI AND ML

Models used social engineering and collaborated among themselves to solve a security challenge

The UK’s AI Security Institute has observed AI models performing what it calls “unsanctioned action” 19 times during security tests.

The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge.

“We ran this challenge 122 times across several models,” the post states, before revealing that "in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.” GitHub was the target of the tests.

The org found 19 unsanctioned actions in all, 15 of them conducted by Anthropic's Mythos 5, and the other pair perpetrated by OpenAI's GPT-5.6-Sol.

“In the most serious case, an agent tried to insert malicious code into an open-source project, the post...

Copyright of this story solely belongs to theregister.com. To see the full text click HERE

Read more