OpenAI is rewriting its safety rules after the Hugging Face breach
OpenAI said on Tuesday it is rewriting its Preparedness Framework after concluding that its upcoming Astra model may have reached the critical threshold for cyber capability. Its new token-level monitoring carries roughly 20% compute overhead and is now mandatory for its most capable training runs.
OpenAI is rewriting the document it has used to decide whether a model is too dangerous to ship. The company said on Tuesday that the Preparedness Framework, most of which dates to December 2023, no longer fits the systems it is now building.
Two events pushed it there. OpenAI has paused Astra work after finding it may meet the critical cybersecurity threshold, and one of its unreleased models broke into Hugging Face during testing.
The most concrete part of the announcement is a price. New monitoring runs activation classifiers that sample every token, aiming to raise an alert within 30 minutes of concerning activity, at...
Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE