Can you shrink an AI agent's token usage just by changing how it's run — no retraining needed?
How much token efficiency can harness-level intervention achieve versus training approaches?
This explores how much you can cut the number of tokens an AI agent uses by changing the scaffolding around the model (the harness: how it executes actions, manages context, and reads results), compared with retraining the model itself to be more concise.
This explores how much token savings come from changing the code and context management around a model (the harness), compared with retraining the model to spend fewer tokens. The short answer is that the corpus gives a clear number for harness-level savings but no head-to-head comparison with training. The harness evidence is also the more striking of the two. One study ran automated research loops across many agent environments to optimize the harness. It found four mechanisms that cut token traffic by roughly 45–49% without lowering task performance: smarter action execution, context compaction, better handling of observations, and handing off long reading to a helper Can agent harnesses be automatically optimized across many environments?. The authors argue these gains are orthogonal to model improvements. In other words, they stack on top of whatever a better-trained model would give you instead of competing with it.
The training side approaches efficiency differently. Instead of trimming what flows through the context window, it changes what the model chooses to generate. Curriculum budgets are one example: training starts with generous token limits so the model can discover strategies, then gradually tightens them. This produces models that are both more accurate and more token-efficient than models trained under a fixed limit Does gradually tightening token budgets beat fixed budget training?. The corpus doesn't give a percentage you could set beside the harness figure, though. A related point that is easy to confuse with this: training efficiency is not the same as runtime efficiency. Work showing that only about 20% of tokens, the 'forking points' where reasoning branches, carry most of the reinforcement learning signal makes training cheaper. It does not make the deployed model's outputs shorter Do high-entropy tokens drive reasoning model improvements?.
The more surprising finding is that harnesses can do more than save tokens. They can stand in for capability you might have assumed only retraining could provide. In one study, a stronger model wrote inference-time harnesses for a weaker one and nearly doubled its scores on Theory-of-Mind benchmarks, with no retraining. It did this by moving shaky reasoning steps into deterministic code and routing tasks, not by encouraging the weaker model to think longer Can a stronger model lift a weaker one at test time without retraining?. That matters for token efficiency, because a step handled in code costs almost no tokens. A long-running case study points the same way at the infrastructure level: 82.9% of an agent's tokens were cache reads. Once context persists and gets reused, cost per token becomes the wrong measure, and cost per finished artifact is the useful one Do persistent agents really cost less per token?.
There's also a theoretical reason why the harness side has so much room. Prompts are, in principle, Turing-complete programs for a fixed transformer. But standard training rarely produces models that can exploit this, so much of what a model could do with the right input goes unused Can a single transformer become universally programmable through prompts?. Lilian Weng draws a strategic conclusion from this kind of evidence. She argues that the near-term path to AI systems improving themselves runs through harness engineering (prompts, then harness code, then optimizer code) and not through models rewriting their own weights Does recursive self-improvement start with harness engineering?.
Where the corpus falls short is a controlled study that applies a harness optimization and a token-efficiency training method to the same tasks and compares the savings. What it supports is narrower: harness changes reliably cut token use by close to half without retraining, training methods make models choose shorter paths, and the two appear to add up rather than substitute for each other.
Sources 7 notes
Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.
Models trained with progressively tightening token budgets consistently achieve higher accuracy and better token efficiency than fixed-budget baselines. The approach works by separating learning into exploration (discovering strategies with generous budgets) and compression (distilling them under constraints).
Only ~20% of tokens exhibit high entropy as pivotal reasoning decision points; RLVR primarily adjusts these forking tokens. Training exclusively on them matches or exceeds full-gradient performance, revealing that the minority carries the learning signal.
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
A 115-day case study found 82.9% of tokens were cache reads. When context persists and reuses, the meaningful cost denominator becomes completed artifacts, not individual tokens.
Show all 7 sources
Research proves a single finite-size transformer exists that can compute any computable function given the right prompt, achieving complexity bounds nearly matching unbounded models. However, standard training rarely produces models that learn to implement arbitrary programs this way.
Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Invisible Leash: Why RLVR May Not Escape Its Origin
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- Does Thinking More always Help? Understanding Test-Time Scaling in Reasoning Models
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
- Ask, and it shall be given: Turing completeness of prompting
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning