Anthropic just shipped a mid-tier model that beats its own flagship at coding—but the efficiency story gets messy when you look past the press release.
The Summary
- Anthropic released Claude Sonnet 5.5 on September 28 at the same list price as Sonnet 5, claiming 30% faster output and up to 30% lower per-task costs
- Sonnet 5.5 tops Opus 5.5 on Terminal-Bench 4.0, Anthropic's coding benchmark, despite being the mid-tier model—and costs half as much per token
- Independent benchmarking found Sonnet 5.5 uses more tokens than any model tested, undercutting Anthropic's cost-per-task claims when pushed to limits
- Context: Opus 5.5 just took the top spot on Text Arena with 1509 points two days before Sonnet's release, setting a new performance ceiling
The Signal
Anthropic is playing a different game than the rest of the foundation model race. While OpenAI and Google chase benchmark points, Anthropic shipped a mid-tier model that outperforms its own flagship on coding tasks for half the price per token. That's not iterative improvement. That's strategic cannibalization of your own product line—the kind of move you make when you're betting on volume over margin.
The company claims Sonnet 5.5 runs 30% faster and costs up to 30% less per task than its predecessor. At first glance, this looks like the efficiency curve everyone's chasing: better, faster, cheaper. But the real story splits when independent testers got their hands on it.
"An independent benchmarking firm found Sonnet 5.5 burns more tokens than any model it has measured."
Here's what's happening: Sonnet 5.5 generates more tokens than competitors for the same task, even as it delivers answers faster. The per-token price is lower, but the token count is higher. For simple queries, you come out ahead. For complex agentic workflows where the model has to reason through multiple steps, those tokens stack up fast.
This matters because the agent economy runs on margin, not speed alone. If you're building an AI agent that runs 10,000 customer service interactions a day, token efficiency is your burn rate. A model that's 30% faster but uses 40% more tokens doesn't save you money—it costs you more while looking cheaper on the spec sheet.
Key tension:
- Anthropic's internal benchmarks focus on per-task cost
- Independent testing reveals per-token efficiency drops under load
- The gap between marketing and reality widens as workloads scale
Opus 5.5's Text Arena victory two days before Sonnet's launch adds context. Anthropic now has the top-ranked model for general reasoning and a mid-tier model that beats it at code. That's an unusual product stack. Most AI labs ladder their models cleanly: good, better, best. Anthropic just broke that ladder.
The implication for builders: you can't just read the headline pricing anymore. You need to measure token consumption in production, not on benchmarks. The cheapest model per token might be the most expensive model per outcome.
The Implication
If you're running agents at scale, instrument your token usage now. The listed price per million tokens is table stakes. What matters is how many tokens your actual workload burns. Sonnet 5.5 might be perfect for quick-hit tasks where latency matters more than total cost. For long-running reasoning chains, you're probably still better off with a model that thinks slower but talks less.
Watch how Anthropic prices Opus versus Sonnet over the next quarter. If they converge, it means Anthropic sees Sonnet as the volume play and Opus as a prestige product. If they diverge further, it means the market is splitting: speed buyers versus efficiency buyers. The agent economy has room for both, but you need to know which game you're playing.