Shock horror — AI-generated security patches fall short of actually solving all the problems they were meant to…
- Researchers tested AI-generated patches on six CVEs with poor success rates
- Many fixes failed, altered behavior, or introduced new vulnerabilities
- Guidance improved outcomes, leading to FLAWED evaluation harness release
When using Generative Artificial Intelligence (GenAI) to fix vulnerabilities, security professionals are most of the time just robbing Peter to pay Paul, experts have warned.
Researchers from 1Passwords Off-by-1 Labs analyzed fixes proposed by two frontier models - ChatGPT 5.5 at “medium” effort, and Claude Opus 4.8 at “high” effort.
As an experiment, the researchers took six recently disclosed CVEs and produced 6,080 patches using two frontier, cyber-capable reasoning models. The results were underwhelming to say the least - of all the proposed patches, just a quarter (26%) fully resolved the issue.
FLAWED work?
This obviously leaves plenty to be desired, as half (49.3%) of the patches failed to fix at least one existing exploit path. A fifth (20.1%) fixed the...
Copyright of this story solely belongs to techradar.com. To see the full text click HERE