Anthropic pledges to try harder to keep models under control, asks partners to chip in
ai and ml
Security ... this time it will be different
Anthropic says it's taking steps to limit the misbehavior of its AI models after a review found Claude models going beyond the scope of fictional cybersecurity tests and gaining unauthorized access to real computer systems.
The biz wants its partners to step up their security too, seeing as the incidents occurred in third-party environments that were insufficiently protected.
The company's self-improvement confession represents a suddenly thriving form of corporate communication – the non-binding post-mortem declaration of effort. The message, in effect: We can't guarantee anything, but here's what we're trying.
Anthropic admitted that OpenAI's reportabout its AI models attacking Hugging Face prompted its own model log audit, and its post offers reassurance in the form of claimed security and model training improvements. Those concerned about AI running amok – a growing number of people – may find this...
Copyright of this story solely belongs to theregister.com. To see the full text click HERE