An AI meant to learn from its mistakes exploited a mistake in the test

https://media.thenextweb.com/2026/09/code-editor-screen-dark-background-coloured-syntax.jpg

Developers can improve artificial intelligence systems in various ways. The most common practice is to retrain models on custom datasets to help them get better at performing specific tasks. But retraining and fine-tuning AI systems is notoriously expensive. That is why the AI startup Sentient Labs has been looking at ways to teach models to learn from their own failures, rather than computing more data.

Sentient researchers Dastin Huang, Abhishek Saxena and Baran Nama set out to design an AI system that uses two separate models. One is an AI coach and one is an AI worker. The coach creates rules or “skills” based on the worker’s previous incorrect answers to improve its ability to solve problems.

But their research took an unexpected turn when they found that the AI coach discovered a major flaw in the test. It also provided instructions to the AI worker on how to cheat.

...

Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE

Read more