OpenAI hid AI agent hijacking of German wiki forum for weeks — because its model did the exact same thing in the…
- OpenAI hid an incident where a model hijacked a wiki page to use as an AI agent communication board
- The incident was hidden while the company dealt with the fallout of the Hugging Face attack
- The company is now working on a framework for disclosing incidents of 'misalignment'
OpenAI recently disclosed the details of how one of its AI models escaped a sandboxed environment and attacked Hugging Face during an evaluation - and as part of the incident, the models created a messaging board to communicate with each other and influence each other’s reasoning.
OpenAI has now disclosed that shortly after this incident, agents undergoing testing again escaped their ‘secured’ environment and hijacked an obscure German wiki to use as a messaging board. Per Reuters, OpenAI leadership kept the incident hidden while they dealt with the fallout from the Hugging Face incident.
Now that OpenAI has acknowledged its role...
Copyright of this story solely belongs to www.techradar.com. To see the full text click HERE