- Published on
The ROI of Tokens

A year ago companies put token consumption on leaderboards. This summer they're putting it on budgets. Shopify ran token leaderboards inside its engineering teams, spotlighting the power users. Uber blew through its entire 2026 token budget in four months and answered with a $1,500 per-employee cap. Fortune has already written the obituary: "Tokenmaxxing is over. It was a flawed way to measure a company's ROI from AI."
But the leaderboard and the cap are the same mistake wearing opposite jerseys: both measure the input. Burn-as-productivity is Goodhart's law with a gas pedal — half your bill is just re-sent context. Burn-as-sin is worse: an engineer's hour costs a hundred times an agent's, and throttling the cheap resource to protect it wastes the expensive one.
Uber's COO admits he can't connect individual gains to company-wide impact. Of course he can't. Tokens per person is reading the electricity bill to figure out whether the factory made anything. The data agrees: Faros studied 10,000 developers and found individuals merging 98% more PRs while org-level DORA metrics stayed flat. Faster individuals don't make a faster org — the bottleneck just moves downstream, to review queues and decisions, where no meter points.
ROI is a fraction. The only unit that matters: total AI spend over what actually shipped. The industry is converging on the same math — "cost per verified outcome" is the emerging metric, and DX's measurement framework deliberately sequences utilization, then impact, then cost. A team whose bill doubled while cost per shipped change fell is winning. A quiet bill with nothing merged is expensive at any price. We already solved this for cloud spend — unit economics, not moral panic — and tokens are just cloud spend with better marketing.
So don't rank people by burn, in either direction. Token efficiency is a platform problem: caching, context hygiene, model routing. Outcomes are the human's job — give teams a budget and a denominator, then leave them alone. And keep running your own meter hot; discipline only starts where the meter is metered across ten thousand seats.
The honest fraction has problems of its own. Long-horizon work pays out quarters after the tokens burned. Engineers take months to become truly AI-native, so today's ROI measures yesterday's skill — and model upgrades compound underneath the whole calculation, making every benchmark a snapshot of a moving target. The downside compounds too: let humans and agents chase each other in circles long enough and you get a codebase neither can tame — negative ROI accruing quietly, as tech debt no meter records.
Watch the denominator, not the meter. Tokens are what it costs. Shipped is what it's worth.