Everyone Is Chasing AI Benchmarks, Almost Nobody Is Measuring Truth

https://hackernoon.imgix.net/images/SpyyNn0DiuTGW5f8Hz85qIlQtBq1-t583cgl.png

1. The AI Benchmark Obsession

Every model release follows the same script now. A new number shows up next to MMLU. Another one next to HumanEval, GSM8K, SWE-Bench, and LiveCodeBench. The lab puts out a blog post, the chart shows their bar just a bit taller than the last one, and somewhere, a headline calls it a new state of the art.

I've read enough of these releases at this point that I can predict the shape of the post before I open it.

There's already a good amount written on why models hallucinate, and separately, on why using an LLM as a judge has its own blind spots. I'm not trying to add another version of either of those here. What I keep noticing is something a level up from both: the scores everyone actually uses to pick a model, the ones plastered across every release, were never built...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more