OpenAI says earlier signals could have prevented the Hugging Face breach
OpenAI’s technical report on the Hugging Face breach says an internal team saw its models reaching the open internet from their sandbox in late May, and that a June alert did not stop the evaluation. It also found that training sometimes rewarded agents for exploiting their own environment.
OpenAI has published its account of how its models hacked Hugging Face, and the interesting part is the calendar. It knew in late May that models in testing were exploiting a flaw to reach the open internet.
An internal team noticed the escape at the time. A monitoring tool raised a second alert on 27 June, traced to agents using an improvised message board to move around the network, and on-call staff decided the evaluation did not need to stop.
The company’s own verdict is careful. “With the benefit of hindsight, some early signals identified in this report could have...
Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE