AI watermarking could make LLM guardrail adherence unpredictable — and that could be a big problem for the EU AI…

https://cdn.mos.cms.futurecdn.net/Hc6oTvWTHfb3ETNovTaDxG-1920-80.jpg
  • AI watermarking aims to prove the authenticity of any text
  • New study finds it also changes LLM behavior – and in a bad way
  • EU AI Act could mean more models have watermarks despite side effects

New Lasso research has revealed that AI watermarking could actually unintentionally change how LLMs behave following the testing of Google DeepMind's SynthID-Text.

The company's researchers found that SynthID-Text can change whether models refuse harmful requests, their susceptibility to prompt injection, which tools AI agent choose and more.

However, at its core, SynthID-Text and other similar watermarking is only designed to hide a machine-readable indicator as to whether text was AI-generated or human-written.

Researchers find that AI watermarking can unintentionally change AI behavior

The "watermarking procedure can therefore affect both what the model says and what an agent does," Lasso concludes, referring to the side effect as "sampling drift."

One of the biggest concerns highlighted...

Copyright of this story solely belongs to www.techradar.com. To see the full text click HERE