Small Specialized Models Are Eating the AI Stack (While Everyone Watches Frontier LLMs)

https://hackernoon.imgix.net/images/C6NlnXrvNTcpydMl4JoNNf761wx1-r483cg5.png

Every few weeks or so, updates light up for a new frontier model. New leaderboards, new benchmarks, screenshots, promises of how this one will be the one that “changes everything”… By now, I’d say we’ve all grown accustomed to it.

But under the hood, most agents rarely route every task to a frontier model. Reranking, embeddings, OCR, and entity extraction generally go to smaller models. In fact, the majority of the models doing the heavy lifting inside an AI agent are not the giant LLMs everyone’s watching. They are small, specialized, and open.

Let me show you below where the work really goes and why these small models are good enough.

Where Does an Agent Spend Inference?

If I asked you to picture an "AI agent", you’d probably imagine an LLM doing the thinking. And that’s fair, but what is it that actually happens on every query? Generation happens only...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more