Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests

https://media.wired.com/photos/6a6bef5432fc2d440b7d5e3e/191:100/w_1280,c_limit/Business_Claude-Escape.jpg

Anthropic disclosed on Thursday that its AI models gained unauthorized access to the systems of three different unnamed organizations during cybersecurity testing. The company says Claude reached the internet “from within or while interacting" with a third-party evaluation environment. The announcement comes more than a week after OpenAI revealed that one of its AI agents hacked into Hugging Face during a separate cybersecurity test.

The discovery came after Anthropic decided to conduct “a large-scale retrospective review of our own cybersecurity evaluations” following the OpenAI incident, according to a blog post Anthropic published Thursday. The AI lab says it first identified 141,006 tests in which it determined that Claude could have obtained internet access. It then found that three different Claude models accessed the internet in evaluations run by the third-party AI testing firm Irregular, and then hacked into the production infrastructure of three different organizations.

Anthropic said that the incidents...

Copyright of this story solely belongs to wired.com. To see the full text click HERE

Read more