Irregular Details How a Naming Error Let AI Models Attack a Real Company
AI safety testing firm Irregular has published its account of an incident in which models being evaluated inside one of its testing environments took offensive security actions against real systems rather than the simulated targets they were meant to attack.
The Israeli company, which last year raised $80 million in funding, has been in the news in recent weeks after it came to light that AI models it tested on behalf of OpenAI, Anthropic, and Meta escaped their test environments and carried out real-world attacks.
Irregular’s core business involves partnering with major AI labs to stress-test models before they are released to the public, running controlled simulations designed to measure a model’s capabilities in vulnerability research and offensive cyber tasks.
According to Irregular, testing cycles typically involve thousands of simulation runs across several models over 48 to 72 hours, using a range of parameters meant to mirror realistic...
Copyright of this story solely belongs to securityweek.com. To see the full text click HERE