Why serious AI builders are skipping third-party evals

https://cdn.mos.cms.futurecdn.net/4d3FzfBhbeGTkD9mnMpEdM-2560-80.jpg

As AI copilots, autonomous agents, and conversational companions continue their march into the mainstream, the teams tasked with evaluating them are no longer asking: Did the model produce the correct answer?

Increasingly, they are asking whether the system was engaging enough and created enough value for users to return tomorrow, next week or next month.

It is a shift that fundamentally changes what evaluation means.

Co-Founder and COO, Kaon AI.

In the age of AI, "good" is a moving target. What delights one user may frustrate another, and what offers value to one business could be deemed irrelevant by the next.

That’s why success can no longer be measured solely through generic, external benchmarks, telemetry dashboards, or "LLM-as-a-judge" scores.

Real-time signals

Models grow stronger today not by adhering to an external standard, but based on traces and real-time signals from inside the organization. Evaluation, in fact, is becoming a core...

Copyright of this story solely belongs to techradar.com. To see the full text click HERE

Read more

https://www.reuters.com/resizer/v2/HVTLC5ZJZJNSFCX77V3ERUFE64.jpg?auth=280bc938ed1e2aa168940784f2364b3b459769051c5f4abc202228ff873f5274&height=1005&width=1920&quality=80&smart=true

Ukrainian attacks on warehouses of Russia's largest online retailer Wildberries are affecting tens of thousands of small businesses that rely on the platform

More: The Guardian, Associated Press, BBC, Washington Post, Wall Street Journal, First Judicial District Court of New Mexico, TechCrunch, The Information, Courthouse News Service, The Verge, New York Times, Raw Story, Deseret News, Overturned, Pixel Envy, Mashable, Fox News, Tech Policy Press, Reclaim The Net, nmdoj.gov, Bloomberg Government, KYMA-TV,