Researchers fear safety disaster ahead of OpenAI’s Astra release

https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/gettyimages-2287521404.jpg?quality=90&strip=all&crop=0%2C10.737892056687%2C100%2C78.524215886627&w=1200

OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it “may be the single worst development for AI security/safety to date.”

Shortly after OpenAI said on Tuesday that it had delayed Astra’s release to work on safety issues, The Informationreported that Astra shows far less of its “thinking” than other frontier AI models, sparking concern it could be dangerously hard to monitor.

Most top AI systems today are built using a technology known as a transformer, which processes some types of information linearly through layers before producing an answer. Models can be made to show their reasoning as they go, essentially “thinking out loud.” This “chain of thought” allows researchers and automated safety systems to monitor what...

Copyright of this story solely belongs to theverge.com. To see the full text click HERE