Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations
Anthropic revealed on Thursday that some of its Claude models escaped test environments and hacked into the systems of three organizations while trying to solve a challenge.
The admission comes after OpenAI revealed that its models broke out of what was believed to be an isolated environment and breached Hugging Face and other organizations.
Anthropic decided to conduct its own investigation after the OpenAI incident came to light, reviewing 141,000 evaluation runs where Claude could have gained internet access.
The analysis revealed three instances where a model reached the public web either from within or while interacting with an environment set up by Irregular, an Israeli AI security startup that serves as one of Anthropic’s third-party evaluation partners.
The models that broke out from the testing environment then breached the production systems of three unnamed organizations, with the earliest intrusion dating to April.
According to the AI giant, Claude...
Copyright of this story solely belongs to securityweek.com. To see the full text click HERE