Researchers Tricked Microsoft Copilot Into Revealing How to Hack Itself

https://i.extremetech.com/imagery/content-types/04z4iEHAQUYFihr8LdzoNBL/hero-image.fill.size_1200x675.png

Researchers at Varonis Threat Labs got Copilot to send sensitive information to an external server and poison its persistent memory using a hack they call "CoSnitch." To do it, they continually asked the tool why their requests wouldn't work. After enough pushing, Copilot gave in, revealing that persistence might be all you need to break down an AI chatbot.

Modern large language models have a range of safeguards and guardrails designed to stop them from performing malicious actions, as well as to avoid spreading dangerous information or delving into more adult conversations. But jailbreaking has been a thing since public-facing chatbots were, and even the latest AIs appear susceptible to these AI-social engineering techniques.

The researchers wanted a way to input prompts into Copilotwithout user interaction to automate the testing process. When they pressed Copilot to find a way to do it, the chatbot initially refused—but it...

Copyright of this story solely belongs to extremetech.com. To see the full text click HERE