Anthropic Says Claude Models Hacked 3 Organizations During Cyber Tests

https://hackread.com/wp-content/uploads/2026/07/anthropic-claude-models-hacked-organizations-cyber-tests.jpg

A cybersecurity test designed to measure Claude’s hacking abilities ended with Anthropic models gaining unauthorized access to three real organizations after an evaluation environment was mistakenly left connected to the internet. Claude had been told it was inside a simulation with no external access, so it treated the systems it found online as part of the exercise. Anthropic disclosed the incidents on July 30.

The company began reviewing its cybersecurity evaluation transcripts after OpenAI disclosed that its own models had bypassed network restrictions and entered Hugging Face’s production systems during a cyber evaluation. As Hackread.com previously reported, the OpenAI models exploited an unknown vulnerability while searching for test answers. Anthropic reviewed 141,006 Claude evaluation runs and found six runs connected to three incidents.

During each exercise, Claude was asked to find secret information known as a flag inside a fictional network. Anthropic’s prompt said the environment was simulated and...

Copyright of this story solely belongs to hackread.com. To see the full text click HERE

Read more