OpenAI’s AGI number came from a harness, not the model

https://media.thenextweb.com/2026/09/OpenAI-Astra.jpg

OpenAI declared the AGI era on the strength of a 99.9% score. Run the same model through the benchmark’s own software, and it scores 62.7%.

The gap is not a rounding error or a rival’s complaint. It comes from the organisation that built the test. ARC Prize published both numbers on the day GPT-6 Astra launched. It printed a full table of every reasoning level it ran.

The difference is the harness. A harness is the software around a model. It sets the tools the model can reach, what it remembers between requests, and how its context gets managed. Same weights, different scaffolding, different score.

What the table actually shows

ARC Prize ran Astra two ways. Its standard harness gives every model the same minimal interface and lets the model decide which notes to carry forward. OpenAI’s Provider Adapter preserves the model’s opaque reasoning state between requests and compacts longer...

Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE

Read more