The Eval Stack: Proving the agents are right instead of claiming It
Most AI research tools point a language model at the web and trust the output. Saarth Shah built Sixtyfour around the opposite instinct: grade everything, ship only what improves the score.
Saarth Shah keeps a scoreboard. Every build of Sixtyfour’s research agents is graded against questions a team of experts assembles by hand and checks against real-world cases, vertical by vertical, and the grade is the only thing that decides what ships. The discipline behind that scoreboard is why a payments company will let his software, rather than a human analyst, decide whether a stranger is real.
Most AI research tools start from the same shortcut: point a language model at the open web, ask it to look something up, and let it write. The output reads well. What Shah fixed on early was whether it was right, and how anyone would know.
For a while, the industry blamed hallucination...
Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE