A New Trick Reveals AI Models’ Inner Thoughts

https://media.wired.com/photos/6a7a6916b0493fdce2877f53/191:100/w_1280,c_limit/Distilling-AI-More-Complicated-Than-You-Think-Business.jpg

Computer scientists recently discovered a way to extract the hidden “thinking” that frontier AI models perform as they work through complex problems.

The findings provide some evidence—although not conclusive proof—that certain Chinese models may have been trained by “distilling” reasoning information from US models that was supposedly hidden because of how closely some of their thinking or reasoning patterns seem to match. The researchers have also demonstrated that the method could be used to recover personal information, like passwords and API keys, from a model’s inner reasoning, although this vulnerability has been fixed.

“All major frontier model providers we tested share this vulnerability,” says Alexander Panfilov,⁩ a computer scientist at University of Tübingen in Germany who was involved with the work. “It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks.”

Panfilov and colleagues from the University of Tubingen, the Max Planck Institute, the...

Copyright of this story solely belongs to wired.com. To see the full text click HERE

Read more