Anthropic spent this week in hot water over cybersecurity

https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/STKS533_AI_AGENTS_HACKING_B.png?quality=90&strip=all&crop=0%2C9.9676601489831%2C100%2C80.064679702034&w=1200

After admitting earlier this year that its AI models had hacked other companies’ systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models’ single-minded “recklessness” — and will likely fuel already raging concerns about cybersecurity and AI.

In Anthropic’s report, it detailed four cases this year in which its own AI models hacked an external company or exploited vulnerabilities. In one, an “internal, general-purpose research model” broke into third-party systems, using access tokens and passwords and downloading files. In another, a Claude model attacked a company with a live web application reachable on the public internet and handled user data. A third model accessed a “machine belonging to a third party that it was able to access” — apparently believing it was part of its evaluation exercise, per Anthropic — then used...

Copyright of this story solely belongs to www.theverge.com. To see the full text click HERE

Read more