Formula Predicts When AI Chatbots Are at Risk of Turning Bad

https://www.securityweek.com/wp-content/uploads/2025/05/AI-Hallucinations.jpg

Researchers from George Washington University have published a paper examining whether the time and cause of AI going rogue can be predicted; and if predicted, prevented.

Since 50% of the world’s population carry devices that can run personal AI companions with no internet connection and limited security, they focused their research here. The lack of cloud-based safety filters, real-time telemetry, live monitoring, or the ability to patch weights once deployed provides a good test bed for analyzing AI’s chat-style transformer behavior when left to its own devices.

If there is a tipping point, it will stem from the AI’s Attention head. This is the computational component that determines which earlier AI tokens are most relevant when processing the current token. The decision lays the foundation for the next token and so on until the chatbot has finished responding to the user’s prompt. The token is loosely, but not precisely, related...

Copyright of this story solely belongs to www.securityweek.com. To see the full text click HERE

Read more