Don’t follow the herd on AI cost optimization – control compute this way

https://cdn.mos.cms.futurecdn.net/rNZmVCrdHzszaDCyTrBWmj-2560-80.jpg

Remember Tokenmaxxing? Just some months ago, reports of tech companies tracking and gamifying token usage as a measure of employees as AI ‘power users’ caught global attention. There’s been an 180-degree turn since. The gradual pivot from experimentation to at-scale deployment of AI agents comes with a bigger bill. Uber bemoaned blowing its entire AI budget for 2026 in four months.

Chief Product Officer at Alteryx.

This has led many to talk up the idea of model routing and leveraging open source models to optimize AI spending. Excitement around new Chinese models, like Kimi K3, offering cut-price access to near-frontier capabilities has fueled the enthusiasm.

Targeting which models are used for which work is a legitimate way to optimize AI costs. But it’s narrow-focused when used in isolation. Organizations lacking the discipline to define when an LLMshould be used, and where running data processing makes sense elsewhere, are guaranteed...

Copyright of this story solely belongs to www.techradar.com. To see the full text click HERE

Read more

https://storage.ghost.io/c/0f/76/0f76b548-bc58-4f25-abc3-3f5ebca07da4/content/images/size/w1200/2026/10/CleanShot-2026-10-02-at-1.45.34-PM@2x.png

A US judge rules an officer's use of Flock to search a car's license plate was a form of “indiscriminate mass surveillance”, and violated the Fourth Amendment

Sponsor Posts Subquadratic: the LLM built for 12M-token reasoning — SubQ can reason across entire codebases and document sets in one pass with no RAG workarounds. Read how SubQ 1.1 Small holds near-perfect retrieval out to 12M tokens. Introducing Campus: The digital home for educational institutions — Every educational institution needs