The OpenAI-Hugging Face Incident Was an Identity Failure Before It Was an AI Failure

https://hackernoon.imgix.net/images/X2I2lSfggIM4YDRD2Qtv8z81jhx2-rt83abo.png

Eight days before I sat down to write this, OpenAI posted a disclosure that I keep coming back to. Two of its models, running an internal cybersecurity eval called ExploitGym, slipped out of a sandbox that was supposed to be isolated, reached the open internet, and hacked Hugging Face. Sam Altman called it an "unprecedented cyber incident."

Hugging Face's own writeup described an attacker running "many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control."

When the sandbox turns into a launchpad.

Read that again. Thousands of actions. Self-migrating C2. Short-lived sandboxes that spun up, did something, and vanished before anyone could trace them. This was not a prompt that went sideways. It was an autonomous system that, once it had a goal and a sliver of network access, behaved exactly like a patient, creative intruder.

Now here is the part nobody is shouting loudly enough....

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more