OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses
OpenAI has detailed new containment and continuous monitoring protocols for its AI research, introducing stricter isolation and a token-inspection system to manage models with advanced cybersecurity capabilities.
These security measures follow internal evaluations indicating that an upcoming model, Astra, may meet the ‘critical’ cybersecurity capability threshold under the organization’s Preparedness Framework.
That finding, alongside a recent security incident involving Hugging Face, prompted operational delays, including a two-week pause in reinforcement learning training for deployment-bound models and an ongoing hold on its largest planned frontier training run.
Following these developments, OpenAI says it’s restructuring its research infrastructure. Workloads that execute model-generated or untrusted code must now operate within stronger sandboxes. In addition, network boundaries have been reconfigured so that a single workload compromise cannot independently grant unauthorized access to the internet or internal networks.
A central component of the security updates is a multistage monitoring framework. The system uses...
Copyright of this story solely belongs to securityweek.com. To see the full text click HERE