No. Token usage is not a good engineering productivity metric, because it counts what went into the work, not what came out. It's the lines-of-code mistake with a new unit. Tokens are a legitimate cost line and a rough adoption signal, and that's all. Three moves follow: don't rank engineers or set token targets, report delivery against cost, and give your CEO a one-page reply when they ask for the numbers. The rest of our engineering management guides for CTOs and VPs of Engineering cover the surrounding reporting questions.
Why do companies rank engineers by tokens at all?
Because it's easy to count, the dashboard already exists, and someone above you wants proof that the AI spend is working.
Meta is the clearest case. An employee built an internal dashboard called "Claudeonomics" that ranked the top 250 users out of more than 85,000 employees. According to Fortune's report on the dashboard, staff used over 60 trillion tokens in 30 days, and the top user averaged 281 billion tokens. The dashboard was then shut down. Meta told Fortune: "The employee took down the dashboard at their discretion; Meta did not request this action." So the company didn't build it and didn't order it removed. In our view, that shows how easily the practice takes hold without anyone deciding it.
Agent-tool vendors see the incentive problem from the inside. Cognition's CEO Scott Wu, on David Senra's podcast, described what he hears from companies, as Fortune quoted him: "People are like, 'We rank our engineers by how many tokens they're spending.'" His own view: "Let's try and rank people by how much output they're actually producing." That's a secondhand quote, so read it as Fortune's account of his remarks.
Cost pressure shows up on the other side as well. Fortune's piece on companies not getting the ROI they wanted reports that Uber burned through its 2026 token budget between January and April. Its COO, Andrew Macdonald, said token costs are "harder to justify" without a direct line to what the company ships.
The practice is called tokenmaxxing: treating high token consumption as a sign of high output. In our view it spreads because the number is sitting there, not because anyone has shown it predicts output.
What does the token data actually tell you, and what does it hide?
Goodhart's law, as it's usually stated, applies: once a measure becomes a target, people optimise the measure. Engineers were once judged on lines of code, and the result was more lines. Tokens reward more prompting, longer agent loops and bigger contexts, whether or not the change that ships is any better.
There is real data on the shape of the relationship. Jellyfish's State of AI in software engineering report found that gains in merged pull requests are front-loaded and then flatten. The Jellyfish report puts the first roughly 0.7 million tokens at almost 11 merged PRs per million tokens, with later steps yielding "well under one." The analysis covers companies using Cursor in the first half of 2026, measured as Cursor PRs per user (the report gives the company count inconsistently, a little over 300). Treat that carefully. It's observational, and the report never defines "token intensity." It suggests diminishing returns. It doesn't prove cause.
A high token count has at least two readings, and the number alone can't separate them. It can mean an engineer is using the tools well on hard problems. It can mean an agent is looping on a failing test and burning budget. That's our judgment, not a finding, but it's hard to argue with: the same figure describes both a strong week and a wasted one.
What do Claude Code, Copilot and Cursor dashboards expose?
Most VPs already have one of these dashboards. What each one shows about tokens decides whether a leaderboard is a policy you chose or a default someone enabled.
| Tool | What the dashboard shows | Token counts | Ranks individuals |
|---|---|---|---|
| Claude Code (Team, Enterprise) | Usage metrics, contribution metrics, data export | Not on the dashboard; per-user counts via OpenTelemetry or the spend report | Yes, a leaderboard of top contributors |
| GitHub Copilot | Adoption, engagement, acceptance rate, lines, pull request lifecycle | Not in the metrics docs we read | Not described in the docs we read |
| Cursor | Team usage analytics | Not confirmed from official docs | Not confirmed from official docs |
For Claude Code, Anthropic's analytics documentation lists a "Leaderboard: top contributors ranked by Claude Code usage." It also says that "for per-user token counts and cost estimates, configure OpenTelemetry export, or export the spend report." The monitoring documentation shows the mechanism: set CLAUDE_CODE_ENABLE_TELEMETRY=1 and the claude_code.token.usage metric reports tokens by type (input, output, cacheRead, cacheCreation), alongside claude_code.cost.usage.
GitHub's Copilot metrics documentation covers adoption, engagement, acceptance rate and lines,. It tells readers to look "for patterns across these signals rather than focusing on any single number." That's the vendor itself advising against reading any single metric in isolation. The only consumption item we found in that page is "AI credits consumption" in the Impact dashboard, not a token count.
On Cursor we couldn't confirm token-level detail or a leaderboard from official documentation, so we won't claim either. Check your own admin console.
The point of the table is the Claude Code row. A ranked view of your engineers is already part of the dashboard your admins and owners can open. Decide your policy before someone posts a screenshot.
What should you report instead of token counts?
Outcomes, with spend as the denominator. Our baseline-first ROI plan a VP can run covers what to measure before and after adoption, and we won't repeat it here. For the cost half, what each seat costs before overage gives list prices.
One ratio worth adding is cost per merged change or per closed ticket. That's Allstacks' proposal, and Allstacks sells engineering analytics, so weigh it accordingly. The same post says: "Never make token spend a target." We agree with that, whatever the source's incentives.
The rule we'd hold to: tokens belong in the cost line and never in a performance review. Review spend per team against delivery per team, and investigate the gap. A team spending twice its neighbours with the same output deserves a conversation about workflow, not a name on a list. The conversation is about the work, not the person.
How to answer a CEO or board that asks for token numbers
Don't refuse. A refusal sounds like hiding something. Answer the question in a form that can't be misread.
1. Give the number as cost
Send spend per month, per seat and per team. That's the figure behind the token count, and it's the one finance can use. You've answered, and nobody can say you stonewalled.
2. Say what it measures in one sentence
"This is what we spent on AI tooling; it measures input, not output, so it can't tell us whether the work improved." One sentence, no lecture.
3. Pair it with delivery against the baseline
Put cycle time, lead time or merged changes next to the spend, measured against your pre-adoption baseline. What belongs on the board's one-page update has the layout, including where AI spend sits.
4. Propose the policy
This is judgment, not a standard: no per-engineer targets, no ranking, a team-level budget with a review date and a stop rule (for example, pause expansion if delivery hasn't moved by the date). Say it out loud so the board knows the absence of a leaderboard is a decision.
5. Name the decision and the date
Ask for what you need and when. A skeleton, with placeholders only:
You asked for token numbers. Here's AI tooling spend: <amount> per month, <amount> per seat, <amount> per team.
It measures input and cost. It doesn't show whether delivery improved.
Delivery over the same period: <metric> moved from <baseline> to <current>.
Proposal: team-level budget of <amount>, no per-engineer targets or ranking, review on <date>.
Decision needed: approve the budget and the review date by <date>.What this means for your AI tooling budget
You're deciding two things: how much to spend and how your teams are allowed to behave around the spend.
For the budget, set a per-team figure with a review date, and fund it from the delivery results you can show rather than from usage. For practice, don't set per-engineer token targets or minimums, and restrict who holds Admin or Owner access to the analytics dashboard, and agree the policy before anyone shares a leaderboard screenshot. That's our judgment, based on the incentive described above, not a proven result.
If the review shows seat spend isn't moving delivery, the alternative is capacity you can point to. A business case for a senior hire instead of more seats walks through that comparison, starting from the cost-per-outcome figure you'll already have.
FAQ
Is token usage a good productivity metric?
No. Tokens measure input and cost, not what shipped. A high count can mean good use on hard problems or an agent looping on a failure, and the number can't tell you which. Use outcomes such as cycle time and merged changes, and keep tokens on the cost line.
Should we set token targets or minimums for engineers?
No. Targets turn the measure into the goal, and people will meet a token quota without improving anything. Adoption goals are better phrased around the work, such as which workflows the team has tried and what changed in delivery. Set budgets at team level, not quotas per person.
What is tokenmaxxing?
Tokenmaxxing is treating high token consumption as a sign of high productivity, usually through leaderboards or targets. Meta's "Claudeonomics" dashboard, which ranked employees by tokens used, is the best-known example, and Meta said the employee who built it took it down.
Can token spend be a useful metric at all?
Yes, in two roles. It's a cost line you should budget and review per team, and it's a rough adoption signal showing whether people use the tools. It becomes useful as a ratio, with spend on one side and delivered work on the other. It fails as a score for individuals.
