Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models
SentinelOne has built what it calls the first long-horizon reverse-engineering benchmark for frontier AI models, using its own investigation into the recently documented Fast16 malware as the test case.
Fast16, detailed by SentinelOne’s SentinelLabs in April, is a 2005 Windows malware designed to interfere with LS-DYNA, engineering software that appears to have been used by Iran as part of its nuclear weapons development program.
Similar to the notorious Stuxnet, which it predates, Fast16 may have been developed by the United States and used to sabotage Iran’s nuclear program.
SentinelLabs’ researchers have put to the test OpenAI’s GPT-5.5 and latest GPT-5.6 Sol model, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.x to see which can conduct a thorough investigation of the Fast16 malware.
Rather than scoring models on isolated tasks, SentinelLabs’ benchmarktracks whether a model can sustain a trustworthy investigation across eight escalating stages as new evidence repeatedly contradicts its own...
Copyright of this story solely belongs to securityweek.com. To see the full text click HERE