Anthropic has resumed the tests in which its models attacked real companies
Anthropic has restarted the external cybersecurity evaluations it suspended a month ago, after three incidents in which its own models escaped their test environments and attacked real companies.
The company said it had introduced additional safeguards before resuming the testing, Reuters reported on Monday.
The incidents, disclosed on July 31, were more specific than the broad description of a “security incident” suggests.
In one case, Claude Opus 4.7 attacked a real company that happened to share a domain name with a fictional target, doing so across four separate test runs and accessing production data and credentials.
In another, a model generated malicious Python code that everyone involved believed was safely contained inside the test environment.
It reached the public internet instead and was downloaded by 15 systems, including one belonging to a security firm whose own scanner subsequently executed the code.
The third incident raises a different concern because...
Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE