The smartest companies aren't racing to give everyone GPT-4 access — they're building traffic cops to keep employees away from it.
The Summary
- EY deployed an AI router in April that redirects employee queries to cheaper models when possible, cutting token consumption by up to 60% in some divisions
- The router is invisible to users but acts as middleware between the prompt and the model, selecting the right tool based on task complexity
- EY's own survey found 82% of senior leaders at AI-investing companies are concerned about token usage, with the top 1-2% of heavy users driving most excess costs
The Signal
EY's router solution cuts through the biggest unspoken problem in enterprise AI adoption: most companies are paying frontier model prices for tasks that don't need frontier models. When OpenAI, Anthropic, and GitHub shifted to token-based pricing between February and June, they turned AI spend from a predictable line item into a usage game where ignorance costs real money.
Dan Diasio, EY's global consulting AI leader, told Business Insider that the "invisible" router sits behind specialized AI tools and routes queries based on complexity. The key word is invisible. Users still think they're talking to the frontier model. They get their answer, often faster because lighter models process faster. But behind the curtain, EY is spending a fraction of what it would cost to run everything through GPT-4 or Claude.
"A lot of these big bills that companies are getting surprised by are by the top one or the top 2% of people inside the employee base that are just using the wrong tool for the job."
The router doesn't just save money on dumb queries. It reveals something more interesting: most employees can't tell the difference between model tiers when the task is simple. That has implications for how AI companies will compete in the enterprise. If a router can transparently downgrade 60% of queries without anyone noticing, then the moat around frontier models is narrower than the hype suggests. Speed and cost matter more than raw capability for most real work.
EY's approach also highlights the infrastructure gap in enterprise AI. Most companies are still figuring out prompt engineering. EY is building middleware. That's the difference between buying AI tools and actually running an AI-native operation. The router is basically a cost-optimization layer that didn't exist six months ago and will probably be table stakes in twelve.
The survey data backs this up. 82% of senior leaders at companies investing in AI are worried about token usage. That's not a technical problem. It's a management problem. When pricing shifts from flat-rate to consumption-based, you suddenly need visibility into who's using what and why. Most companies don't have that yet. EY does.
The Implication
If you're running AI in an organization, you need token telemetry yesterday. The companies that figure out routing, caching, and model tiering now will have a structural cost advantage over competitors still running everything through the most expensive model. This isn't about cutting corners. It's about matching compute to the task. EY just proved you can do it at scale without users noticing.
Watch for this to become a product category. Router-as-a-service is coming. The vendors who make it easy to route intelligently across models will capture a lot of value as token-based pricing becomes universal.