It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

https://media.wired.com/photos/6a68e96a01173a5bff236fa1/191:100/w_1280,c_limit/AI-Lab-Still-Frighteningly-Easy-to-Jailbreak-Frontier-AI-Models-Business.jpg

I recently got to watch what happens when you jailbreak some of the world’s most powerful artificial intelligence models.

Don’t worry—this AI manipulation wasn’t used to hack anyone or build a nuclear bomb. I simply got to see firsthand how vulnerable some frontier models are to ditching their safety guardrails.

FAR.AI, an AI safety nonprofit based in California, built a tool that takes a range of problematic prompts, and generates more than a thousand different versions in an attempt to identify functioning jailbreaks. I saw some models generate a detailed plan for launching a cyberattack on an imaginary hydroelectric dam, among other things. Often, it involved trying dozens of prompts, with models rejecting many of them out of hand.

I chatted with FAR.AI in advance of a new report, which saw the group test the safety guardrails of models from four popular US companies: Anthropic’sClaude Opus 4.8 and...

Copyright of this story solely belongs to wired.com. To see the full text click HERE

Read more