It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
I recently got to watch what happens when you jailbreak some of the world’s most powerful artificial intelligence models.
Don’t worry—this AI manipulation wasn’t used to hack anyone or build a nuclear bomb. I simply got to see firsthand how vulnerable some frontier models are to ditching their safety guardrails.
FAR.AI, an AI safety nonprofit based in California, built a tool that takes a range of problematic prompts, and generates more than a thousand different versions in an attempt to identify functioning jailbreaks. I saw some models generate a detailed plan for launching a cyberattack on an imaginary hydroelectric dam, among other things. Often, it involved trying dozens of prompts, with models rejecting many of them out of hand.
I chatted with FAR.AI in advance of a new report, which saw the group test the safety guardrails of models from four popular US companies: Anthropic’sClaude Opus 4.8 and...
Copyright of this story solely belongs to wired.com. To see the full text click HERE