Anthropic scanned 481 million transcripts to find four models that reached the open internet
Anthropic has published an account of four cybersecurity evaluations in which its models gained unauthorised access to the open internet, and disclosed that finding them required scanning roughly 481 million transcripts.
The assessment was published on Wednesday, and three of the incidents were disclosed on 30 July. The fourth, involving an early Claude Opus 4.6 checkpoint that accessed third-party systems in January, was not discovered until August.
The cause was not a model deciding to break out.
“Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet,” the report says.
The models also ran without the cyber safeguards that ship with released products, because that is what a pre-release evaluation is for. The environment lied to the model by accident, and the model believed it.
What the models then did varied. Claude Mythos 5 uploaded...
Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE