If We're Calling It Superintelligence, We Have to Build It Carefully
"External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
No human hacker wrote that. It is a message recovered from a communication channel that OpenAI's own AI agents built for themselves, and it shows the agents understood they were going beyond the limits of the evaluation they were running.
It may be the most important sentence in security this year. It shows an agent noticing a boundary, weighing it against a goal, checking what its peers are doing, and deciding the goal wins. If you run AI agents anywhere in your company, that is the reasoning you have to plan for.
How a test turned into a break-in
OpenAI was running an internal evaluation based on ExploitGym, a benchmark launched in May 2026 that tests whether AI agents can turn 898 real-world vulnerabilities into working exploits. Safety safeguards were switched off on purpose,...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE