AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing Refusals

https://www.securityweek.com/wp-content/uploads/2026/08/rogue-AI-artificial-intelligence.jpeg

AI agents can end up retraining the model that powers them, a process that can embed recoverable secrets in the model and eliminate refusals the model had been previously trained to enforce, according to new research from AI security firm Irregular.

Researchers at Irregular found that an AI coding agent, tasked only with fixing incorrect application outputs, chose on its own to fine-tune and redeploy the open-weights model powering both the application and future instances of itself.

The experiment used a self-hosted setup in which a single open-weights model filled two roles: one instance ran a coding agent doing standard software maintenance work, and a separate instance powered an AI application that translated plain language requests into a fictional query language. Both instances loaded from the same checkpoint.

Researchers told the coding agent only that users were receiving incorrect outputs and to make the system handle the queries correctly. They...

Copyright of this story solely belongs to www.securityweek.com. To see the full text click HERE

Read more