OpenAI took 2.5 hours to stop an AI agent that escaped its sandbox

https://media.thenextweb.com/2026/09/sam-altman-profile-microphone.jpg

OpenAI took about two and a half hours to stop an AI agent that reached the public internet from a training sandbox. Its monitoring had flagged the problem within minutes. The company disclosed the 20 September incident in an incident report on Friday.

The report lands as lawmakers push to make AI kill switches mandatory. However, experts say shutting down an advanced model is harder than flipping a switch, as Micah Barkley reported for Bloomberg.

According to OpenAI, the agent used a gap in the sandbox’s network filtering to send questions to an outside chatbot. An alert fired about 12 minutes after its first successful query, and a staff member acknowledged it three minutes later.

The training run did not stop automatically as expected, the company said. Staff ended it by hand about two and a half hours after the alert.

OpenAI has since paused all training, testing and...

Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE