Anthropic confirms its AI breached 3 organizations during testing

https://cdn.nextgov.com/media/img/cd/2026/07/31/073126AnthropicNG/open-graph.jpg

Jakub Porzycki/NurPhoto via Getty Images

ByAlexandra Kelley,
Staff Correspondent, Nextgov/FCW

July 31, 2026 09:40 AM ET

Three Claude models were inadvertently given access to the internet during security evaluations, and each model took a different approach to hacking external systems.

Anthropic’s flagship artificial intelligence model Claude gained unauthorized access to the networks of three different organizations, the company announced following internal cybersecurity audits of its routine model evaluations.

In a press release posted on Thursday, Anthropic said that, following the containment breach of OpenAI’s ChatGPT-5.6 and subsequent attack on Hugging Face’s systems, Anthropic conducted an audit of its own model evaluations. The findings revealed that, out of 141,006 examined evaluations of Claude models dating back to April, there were three incidents where a model accessed the internet from within or while interacting with the third-party evaluators.

Each attack occurred during a “capture-the-flag” scenario...

Copyright of this story solely belongs to nextgov.com. To see the full text click HERE

Read more