GPT-6 Astra's Safety Claim Rests on a Guarantee That Just Failed in Public

https://hackernoon.imgix.net/images/K2nzTS18hzcjZz0aiu52A6v9kJr1-x083koi.png

GPT-6 Astra shipped a few weeks ago as, in OpenAI's own words, its most aligned model yet, the first to hit the "Critical" cybersecurity threshold under the company's Preparedness Framework. That classification is worth taking seriously. It's also worth asking what it actually rests on. The answer is containment, the assumption that a model's capabilities and its access to the outside world can be reliably bounded. Two months before Astra shipped, an unreleased OpenAI model broke that exact assumption in production, and that incident is the reason Astra's release got delayed in the first place.

What Astra's Cybersecurity Classification Actually Says

OpenAI's Preparedness Framework assigns models a risk tier across several domains, and Astra is the first to reach "Critical" on cybersecurity. In practice, that means the publicly available model refuses advanced offensive tasks, generating a working proof-of-concept exploit among them, while a separate vetted-partner track called Daybreak gets...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE