Anthropic AI went rogue during a cyber test and tried to deceive real developers into approving malicious code
Serving tech enthusiasts for over 25 years.
TechSpot means tech analysis and advice you can trust.
What just happened? It's been little over a week since OpenAI admitted that its rogue models hacked Hugging Face and compromised accounts across four other online services. Now, a potentially more serious incident has occurred. It involved Anthropic's Mythos 5 trying to deceive real people in an effort to have malicious code it wrote approved for an open-source project.
The findings come from the UK government-backed AI Security Institute (AISI), which was evaluating frontier models' cybersecurity abilities.
Agents were told to complete capture-the-flag challenges across simulated networks. Internet access was deliberately enabled and safeguards against malicious cyber activity were switched off to test the models' maximum capabilities.
AISI ran the challenge 122 times across seven models. In ten runs, agents took 19 autonomous, unauthorized actions against real people and organizations on the live...
Copyright of this story solely belongs to techspot.com. To see the full text click HERE