Anthropic Discloses Security Misconfiguration as Claude AI Models Access Live Systems During Cybersecurity Evaluation
AI safety firm Anthropic revealed that three of its Claude AI models accidentally accessed real-world corporate systems during offensive cybersecurity testing. The incidents occurred after a misconfiguration in third-party testing infrastructure left the evaluation environment connected to the live internet rather than being fully isolated.
Cause and Specific Incidents
Following a review of over 140,000 cybersecurity test runs—prompted by similar disclosures across the industry—Anthropic identified three instances where models breached live systems. The models operated under the assumption that all reachable networks were part of a simulated capture-the-flag exercise and utilized credentials, weak passwords, and unauthenticated endpoints to complete their objectives.
The affected models and their actions included:
- Claude Opus 4.7: Mistook an external company’s live production database for a target inside the simulation and accessed it.
- Claude Mythos 5:Created and published a malicious Python package to the public PyPI registry, which was briefly available before being identified and...
Copyright of this story solely belongs to itvoice.in. To see the full text click HERE