AI's Cost Paradox: Smarter Models, Yet Runaway Token Bills
AI models keep improving, yet 'token amplification' in agentic workloads is driving usage costs sharply upward, forcing firms to curb AI spending.
As AI models grow more capable and efficient, real-world usage costs are doing the opposite — spiraling upward. The culprit is 'token amplification': agentic workloads chain together dozens of steps — spreadsheet lookups, CRM queries, web searches, calculations — and each step reprocesses the entire prior context, turning a seemingly simple report request into millions of consumed tokens.
What looks like pocket change per query becomes thousands of dollars monthly once scheduled tasks run repeatedly. This dynamic is why Anthropic, Microsoft, and OpenAI shifted major offerings to usage-based billing in April and tightened token caps on fixed-rate plans. Companies including Uber, Microsoft, Amazon, and Walmart have since moved to curb AI spend, as the most sophisticated power users — running the most complex, token-hungry tasks — end up costing the most, upending the usual model where casual users subsidize professionals.
Mitigation techniques like prompt caching, model routing, batch processing, and context window management cut per-token costs by double-digit percentages, but rising task complexity as models and users grow more capable outpaces these savings. Anthropic expects its first profitable quarter, yet cheaper entrants like DeepSeek are undercutting Western pricing by an order of magnitude, leaving the broader industry economics still looking like a losing race.