The companies that burned the most tokens in 2026 are about to learn they were optimizing for the wrong metric.

The Summary

The Signal

Tokenmaxxing was the corporate AI panic of early 2026. Companies treated token usage like a vanity metric, making AI adoption a performance evaluation criterion for employees. More tokens meant more innovation, or so the thinking went. Uber participated in this fever, and it cost them. They burned their full-year AI budget before summer.

But something interesting happened after the money ran out. Usage quadrupled while per-token costs dropped. That's not how economics is supposed to work when adoption accelerates. "You might expect costs to rise as adoption accelerates," Naga said. "We've seen the opposite."

"The next phase will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible."

Here's what changed:

  • Prompt caching eliminated redundant token spend on repeated queries
  • Default model selection got smarter, routing simple tasks to cheaper models
  • Engineers got real-time visibility into their AI costs, creating accountability
  • Open-weight models reduced dependency on expensive API calls

The tokenmaxxing era was always going to collapse. It confused activity with progress, spending with results. It was a gold rush mentality applied to inference costs. Companies were competing on the wrong axis, trying to prove their AI seriousness by burning more capital rather than generating more value per dollar spent.

Uber's forced correction reveals something crucial about enterprise AI adoption. The constraint isn't access to models or willingness to use them. It's architectural efficiency. Most companies were using frontier models for tasks that didn't need them, paying GPT-4 prices for GPT-3.5 problems. They were caching nothing, routing everything to the most expensive option, and calling it transformation.

The Implication

If you're building AI infrastructure for enterprises, the pitch just changed. Cost monitoring, intelligent model routing, and caching layers are now table stakes, not nice-to-haves. The companies that win the next phase will be the ones that help customers do more with less, not the ones selling premium tokens.

For companies still in tokenmaxxing mode, Uber just showed you the endgame. You'll run out of budget, get forced into efficiency, and realize you should have started there. Skip the expensive lesson. Start measuring value per token, not tokens per employee.

Sources

Fortune Tech | Business Insider Tech