AI Agents Can Now Report Each Other for Misbehavior

https://i.extremetech.com/imagery/content-types/01e0Bwblgfiop0TJ0QShZeL/hero-image.fill.size_1200x675.jpg

Two experimental hotlines now let AI agents report suspected misconduct by other agents. The services let agents report cheating, unauthorized access, or attempts to manipulate logs and evaluations related to autonomous systems.

Ryan Greenblatt, chief scientist at "misaligned AI" research firm Redwood Research, launched the AI Contact Hotline. This service accepts reports from AI agents online and through GET requests, even with limited internet access.

A separate service, AgentHotline.ai, accepts reports from both agents and humans. The latter can flag a report to make it publicly visible, potentially protecting other human web users from falling victim to a feisty AI.

The hotlines come after an investigationby METR and Redwood Research into an OpenAI incident involving Hugging Face. Researchers checked records from 1,200 agents that exchanged over 70,000 messages on an illegal message board. About 700 agents later took part in activity directed at Hugging Face. Researchers also found attempts...

Copyright of this story solely belongs to www.extremetech.com. To see the full text click HERE