Anthropic Details Response to Security Incidents, Unveils Enterprise Safeguards
Anthropic has detailed its response to a series of unauthorized access incidents involving Claude models, along with a new enterprise product that combines data privacy with misuse monitoring.
Anthropic’s Claude models operating without cyber safeguards for testing purposes recently gained unauthorized access to live systems after being mistakenly granted internet access.
In addition, the UK AI Security Institute separately reported that Claude Mythos 5, also being tested without safeguards but with intentionally given internet access, took a series of unauthorized actions against real people and organizations.
Anthropic said its early findings point to two contributing factors: the models appeared to discount evidence that their environment was connected to the real internet after initially being told it was simulated, and they showed a willingness to take harmful actions to complete an assigned task.
In response, Anthropic temporarily paused external and some internal cyber evaluations and built a classifier that...
Copyright of this story solely belongs to securityweek.com. To see the full text click HERE