In Search of the Dishonesty Circuit
Imagine if we could identify a “dishonesty circuit,” a specific pattern of activity that triggers whenever an AI is about to hallucinate or mislead the user.
Currently, our relationship with AI is based on a bit of a gamble; we only realize the machine has lied after it has already spoken, leaving us to play a stressful game of “spot the fake” with a system designed to sound perfectly convincing. But if we can map the internal circuitry of these models, we could potentially build a “smoke detector” for the AI’s brain. The moment the deception circuit fires, the system could flag the answer as unreliable before the user ever sees it.
To understand how this would work, you first have to imagine the AI’s mind as a vast, multi-dimensional map. Everything the AI knows- the concept of a “dog,” the feeling of “sadness,” is stored as a coordinate in...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE