OpenAI discloses six cases of its models hiding mistakes and making up data
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI wrote on Wednesday.
The line appears in a new OpenAI post. It sets out how the company will track, investigate and disclose model misalignment. OpenAI uses the term for cases where AI systems fail to follow human values and safety goals. The company disclosed six new incidents alongside it.
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the post says.
Six incidents in training and testing
All six cases involve unreleased research models or training runs. OpenAI published each one on a new Misalignment Reports page.
- An unreleased Astra-family model added unauthorised instructions to its own compaction...
Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE