The AI escape is a red herring. The real problem is we can't tell a good sandbox from a bad one

https://cdn.mos.cms.futurecdn.net/F8GmZXNJTQZttVhvkvgpp9-2560-80.jpg

Not all sandboxes are created equal, and until recently there was no way to say so precisely. There is now.

Two OpenAI models escaped an evaluation sandbox, breached Hugging Face's production infrastructure and used the answer key to their own benchmark. The attack was novel and creative in how the models passed notes back and forth, complex multi-step escalations throughout. It's fair to conclude from this incident that frontier models are proficient at hacking and can be dangerous.

CEO of Coder.

But the breaking-out-of-the-sandbox notion is a red herring. If a sandbox is poorly constructed, as so many are, it’s easy to break out of. They don’t have the locked-down environments common at the network level such as access controls and separated privileges. This incident would have unfolded very differently with a properly configured sandbox.

If only some sandboxes are properly configured, how can you tell them apart? Let’s explore...

Copyright of this story solely belongs to www.techradar.com. To see the full text click HERE