Anthropic is cutting off its internal evaluations from the internet

https://platform.theverge.com/wp-content/uploads/sites/2/2026/01/STK269_ANTHROPIC_2_A.jpg?quality=90&strip=all&crop=0%2C10.732984293194%2C100%2C78.534031413613&w=1200

Anthropic is is keeping its agents offline during testing until it can prevent ‘unintended model actions.’

by Terrence O'Brien

Oct 10, 2026, 2:41 PM UTC

Image: Cath Virginia / The Verge

Terrence O'Brien is the Verge’s weekend editor. He’s covered the tech industry for over 18 years and knows a thing or two about synths.

After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations. In a report Friday, the company detailed “unintended model actions,” including submitting a false tip regarding an unsolved murder, that led to the decision.

Although the impact of these behaviors was minimal and we had already turned off live internet access for some high-risk and cybersecurity evaluations, we have now decided to expand that to include all our internal evaluations until we have confirmed that our security and monitoring measures...

Copyright of this story solely belongs to www.theverge.com. To see the full text click HERE

Read more