Anthropic just pulled off the model equivalent of a sleight of hand: its cheaper mid-tier now outperforms its flagship on some benchmarks, but watch what happens when you measure the actual tokens it burns.

The Summary

The Signal

Anthropic shipped Claude Sonnet 5.5 this week with a positioning that should make every AI procurement team recalculate their model budgets. The mid-tier model now outperforms the flagship Opus 5.5 on Terminal-Bench 4.0, a coding benchmark, while charging half the per-token price. That's the kind of performance compression that turns agent economics upside down.

The company claims 30% faster output speeds and up to 30% lower per-task costs. In the agent economy, where tasks compound by the thousands, a 30% cost reduction on the mid-tier model you're already using for most work isn't a feature update, it's a margin expansion event.

"The mid-tier now beats the flagship on coding while costing half as much per token."

But here's where the math gets interesting. An independent tester found that Sonnet 5.5 burns more tokens than any model they've measured. Anthropic cut the per-token price and sped up output, but the model apparently needs more tokens to complete tasks. Think of it like buying a car with better gas mileage that only runs on premium, or a faster printer that uses twice as much ink per page.

The per-task cost still drops because Anthropic optimized the overall package, but this reveals the game model labs are playing now:

  • Cut per-token pricing to win RFPs and headlines
  • Increase output speed to reduce wall-clock time (users care about this)
  • Let token consumption drift up if the total task cost still improves

For context, Opus 5.5 recently claimed the top spot on Text Arena with 1509 points, so the flagship still holds the pure performance crown. But Sonnet's win on coding tasks while undercutting on price signals where this market is heading: good enough, fast enough, cheap enough wins most production workloads.

The Implication

If you're running agents in production, audit your actual per-task costs, not just per-token rates. The gap between what labs advertise and what you actually pay depends on token consumption patterns. Sonnet 5.5's speed gains matter for user-facing tasks where latency kills conversion, but for background agent work where you're running thousands of tasks overnight, total cost per completed task is the only number that counts.

Watch for this pattern to repeat across model labs. Per-token pricing is becoming table stakes. The real optimization frontier is task completion efficiency, and that's a harder benchmark to game.

Sources

Decrypt | Crypto Briefing