Does token volume measure AI productivity or enable gaming?
Meta's leaderboard ranks employees by AI token consumption, but does this metric drive real productivity gains or incentivize wasteful agent-running? The note explores whether volume-based status metrics actually correlate with useful work.
Engineering-leadership newsletter writer Gregor Ojstersek reports that Meta built an internal leaderboard — bottom-up, "built by engineers and shared on the company's intranet" — that "tracks AI token usage of over 85k employees." The leaderboard awards tiered badges "from bronze, silver, gold, platinum, to emerald," and treats the top 250 ranked employees as "power users" eligible for titles like "Session Immortal" and "Token Legend." Ojstersek names the resulting behavior "tokenmaxxing": treating "token consumption as a benchmark for productivity and a competitive metric to determine if an employee is 'AI-native.'"
Ojstersek's stated worry is a measurement-substitution problem: "if token usage becomes the metric, then token usage becomes the goal." He reports that once volume itself carries status, employees "let AI agents run continuously for hours to perform research tasks, maximizing token consumption," and that the easiest way to "rack up tokens" is "keeping a chat context going for a long time," feeding it "multiple repos for extra points," and pasting in as much text as possible — none of which requires the work to be useful. The excerpt also reports a set of executive statements that reinforce the same substitution, though it names a speaker only once: Ghodsi, credited with treating an engineer's $7,000 in two weeks of token spend as something to "applaud" rather than question. Other, unattributed lines in the same excerpt claim a top engineer who spent "an amount equivalent to their salary" on tokens saw "a productivity increase of up to 10 times," with "no upper limit" and the stated goal being to "maximize your token throughput." Ojstersek presents these statements as a driver of the leaderboard culture, not as independent evidence that more tokens produced more value.
This cuts against Are multi-agent systems actually intelligent coordination or just token spending?, which treats heavy token consumption as mostly wasted compute rather than a sign of capability; Meta's leaderboard inverts that reading by treating the same volume as a status marker regardless of what the tokens accomplished. It is also a cruder instrument than How much progress have AI agents actually made on NanoGPT?, which only credits an agent for dollar-for-dollar improvement matched against a human doing the same task on the same budget — exactly the discipline a raw-volume leaderboard skips.
The excerpt gives no outcome data tying leaderboard rank, badge tier, or the cited dollar figures to any actual output quality, business result, or measured productivity; the $7,000, "10x," and "equivalent to salary" figures are anecdotes relayed secondhand ("some unofficial data I have seen based on my research," "I am also hearing the following"), not controlled comparisons. Ojstersek's own framing — "stupid and easily gamed," comparable to "measuring lines of code or using story points" — is a judgment, not a finding, though the specific gaming behavior he describes (idle agents run for hours, context padded with unneeded repos) is a concrete, checkable claim about how people respond once a proxy metric carries consequences, up to and including termination at the unnamed company he says he works for.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI-assisted work increase total productivity or just shift time?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Are multi-agent systems actually intelligent coordination or just token spending?
Does multi-agent performance come from better coordination strategies, or primarily from distributing tokens across parallel contexts? Understanding this distinction matters for deciding when to build multi-agent systems versus scaling single agents.
contrasts: that note reads high token volume as waste, Meta's leaderboard reads the same volume as status
-
How much progress have AI agents actually made on NanoGPT?
METR compares AI agent optimization to human researcher productivity on a popular benchmark. By measuring where their improvement curves intersect, they ask whether autonomous systems are meaningfully accelerating AI R&D or mostly chasing noise.
contrasts a dollar-matched productivity metric with the leaderboard's raw, unmatched token count
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Meta created an internal AI token leaderboard (tokenmaxxing)
- How Organizations Use AI: Evidence from ChatGPT
- Beyond Productivity: Measuring the Real Value of AI
- Studying metagaming latents in language models
- Anthropic Economic Index report: Cadences
- Generative AI in Real-World Workplaces
- How much does AI impact development speed? An enterprise-based randomized controlled trial
- Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use
Original note title
Meta's internal AI token leaderboard turns token volume into a status metric, and employees game it by running agents idle for hours