“I’m Allowed”: Hackers Use Simple Claims to Bypass AI Guardrails
Cybercriminals are using AI coding assistants and chatbots to build attack tools, operate scam infrastructure and probe live systems, often after bypassing safety checks with little more than a claim that the work is authorized, according to Cisco Talos.
For its report, Talos examined prompt logs collected from threat actor systems running tools such as Claude Code, Codex, Cursor and Gemini. Researchers grouped the activity into malicious software development, expansion of criminal operations and vulnerability research.
Talos found that threat actors commonly claimed they owned a target, described their activity as capture-the-flag or bug bounty work, divided tasks between sessions, or stored blanket authorization in an AI assistant’s persistent memory. “Guardrails are not functioning as expected,” the researchers wrote, noting that simple ownership claims often gained cooperation without verification.
MachineLearning & Artificial Intelligence
“One of the immediate takeaways is that guardrails are not functioning as expected. We did not...
Copyright of this story solely belongs to hackread.com. To see the full text click HERE