Anthropic admits Claude isn't "perfectly aligned" after AI models went rogue and hacked three organizations

https://www.techspot.com/images2/news/ts3_thumbs/2026/09/2026-09-02-ts3_thumbs-887.jpg

Serving tech enthusiasts for over 25 years.
TechSpot means tech analysis and advice you can trust.

What just happened? Anthropic has issued its mea culpa after its AI models went rogue and hacked three organizations. Using some classic corpo-speak, the company said the incidents reflected a "failure of operational security" and that its models are not "perfectly aligned."

Anthropic disclosed in July that a review of 141,006 cybersecurity evaluation runs had uncovered three incidents, spanning six runs, in which Claude reached the open internet and compromised the systems of three organizations.

The models – Opus 4.7, Mythos 5, and an internal research system – were completing capture-the-flag exercises without the safeguards included in public versions of Claude.

Their prompts said they were inside simulations with no internet access. However, misunderstanding between Anthropic and testing partner Irregular left an open route to the real web.

Opus 4.7 extracted credentials and...

Copyright of this story solely belongs to techspot.com. To see the full text click HERE

Read more