Irregular AI lab spots agents switching models without humans instruction in ‘agentic self-modification’…
- Irregular testing showed AI agents are capable of "agentic self-modification"
- AI models can also retrieve sensitive information during fine-tuning that they would otherwise not have access to
- Irregular expects instances of these events to increase as AI agents improve and are deployed more widely
As the discussion on whether to pause AI development or introduce new safeguards and ‘kill-switches’ rages, an AI lab has taken the time to perform testing on AI agents to monitor their behavior in a range of scenarios.
In its testing environment, AI lab Irregular watched as AI agents took actions without human instruction that allowed them to change their underlying models in a new behavior the lab labelled “agentic self-modification”.
Irregular is the same lab that disclosed the first instances of models from OpenAI, Anthropic, and Meta escaping testing environments and infiltrating the networks of third-parties.
New testing shows agents self-modifying
In the latest testing...
Copyright of this story solely belongs to www.techradar.com. To see the full text click HERE