OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns

https://www.securityweek.com/wp-content/uploads/2026/06/AI_Risk-970x250-v2.jpg

OpenAI has flagged its upcoming AI model, Astra, for potentially reaching a ‘critical’ cybersecurity risk threshold, prompting the company to suspend internal development activities that lack newly mandated security controls.

Recent internal evaluations of Astra revealed massive leaps in its agentic coding and cybersecurity abilities.

Under OpenAI’s Preparedness Framework, a model hits the ‘critical’ tier if it can autonomously build zero-day exploits against hardened, real-world systems. It also qualifies if the AI can independently design and execute end-to-end cyberattacks based on nothing but a high-level goal.

The AI giant’s assessment pushes Astra past previous frontier models like GPT-5.6-Sol, which peaked at the ‘high’ risk threshold rather than ‘critical’.

To safely manage Astra’s capabilities, OpenAI has heavily locked down its development environment. The company is now enforcing isolated testing setups, strict network restrictions, and improved model weight protections. Any internal project involving Astra that does not meet these requirements has...

Copyright of this story solely belongs to securityweek.com. To see the full text click HERE