The Token Price Fallacy: Why Your Agentic AI Bill Keeps Growing While Unit Costs Collapse

https://hackernoon.imgix.net/images/RTmhtsJkZXOhbA3GffJTKGDowAZ2-s483etz.jpeg

Cost per token is falling fast. Cost per completed task isn't, and almost nobody is measuring the difference.

The cost to query a model at GPT-3.5-level performance fell 280 times between November 2022 and October 2024, from $20 to $0.07 per million tokens, according to Stanford HAI's 2025 AI Index Report. By every unit-price chart, inference should be close to free by now.

McKinsey's July 2026 research on agentic AI spend found the opposite in practice: 93% of enterprises running agentic AI are exceeding their AI budgets, and one in five has constrained its use due to cost. Enterprise LLM spending tripled over the 12 months to the end of 2025, while unit prices kept dropping over the same period.

I research token optimization and CO2-aware inference alongside my day job leading Salesforce CPQ and CLM work at one of the largest telecom companies because I keep running into...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more