Do base models find more solutions than post-trained ones?
Explores whether post-training sacrifices solution diversity for single-shot accuracy in agentic tasks, and whether base models with simple prompting can recover that lost coverage at scale.
The paper extends the RL "sharpening hypothesis" — that post-training amplifies a few high-reward behaviors already present in the base model, trading solution coverage (pass@K) for single-shot accuracy (pass@1) — from math and coding into agentic tool-use tasks, a domain where the authors expected this trade-off not to hold because "multi-turn tool use and interaction may require capabilities newly acquired during post-training." Testing 14 base/post-trained checkpoint pairs across four model families (Gemma-4, Ministral-3, Qwen2.5, Qwen3.5) on three agentic benchmarks (BFCL v4 multi-turn, WebShop, ACEBench — 42 cases total), they find base models equipped with only "a simple system prompt with relaxed tool-calling and parsing interfaces," no fine-tuning, "catch up to their post-trained counterparts as the rollout budget K grows" and eventually surpass them: on WebShop, a harnessed gemma-4-31B base model reaches "over 85% pass@128... vs. 56% for RL."
The mechanism they give is bimodalization: post-training does not uniformly raise or lower per-task success probability, it pushes the distribution toward two extremes — "always solved or never solved" — which raises pass@1 and makes repeated rollouts on the same prompt agree with each other (high "pass^K" consistency), but forecloses the long tail of rare-but-reachable solutions the base model can still find at large K. They formalize the resulting gap as "Sharpening Tax," Tax_X(K) = X_Base(K) − X_Post(K), a scalar estimable from a handful of rollouts and reported as "prevalent in most settings" across the 42 model-benchmark pairs. A second finding qualifies the first: the rollout budget k* at which base models cross over post-trained ones shrinks as model scale grows, so whether sharpening helps or hurts must be judged jointly with scale and test-time budget. To offset the tax during training, they propose posterior-tempered group sampling (PTGS), which adapts sampling temperature to per-prompt difficulty and "pays a smaller tax than the fixed-temperature baseline."
This extends, rather than repeats, Do base models already contain hidden reasoning ability? and Does RL teach reasoning or just when to use it?: both describe post-training as eliciting or scheduling capability the base model already holds rather than installing new capability, and this paper supplies the first coverage-based evidence for that claim outside math and coding — in multi-turn tool use, where the authors themselves expected the opposite going in. It sits in tension with Can reinforcement learning discover reasoning strategies base models cannot?, which reports RL-trained models beating base models at all pass@k levels; the difference is training regime — that result comes from prolonged RL with explicit entropy control and reference-policy resetting against collapse, while Sharpening Tax is measured on standard open-source post-trained checkpoints whose training recipes are undisclosed, so the two findings may mark different points on the same entropy-collapse spectrum rather than a contradiction. The bimodalization mechanism also gives a concrete behavioral signature to the collapse named in Does policy entropy collapse limit reasoning performance in RL?: rather than entropy declining smoothly, per-task success rates polarize toward 0 or 1.
What the excerpt does not establish is whether the tax is intrinsic to RL post-training itself or an artifact of the specific open-source checkpoints studied — the authors note their analysis is "observational rather than interventional" since the exact pre-training and post-training data and recipes are undisclosed. It also leaves open whether frontier-lab pipelines already using diversity-preserving techniques (of the kind PTGS or prolonged RL represent) pay the same tax as these academic checkpoint pairs. The stated implication, at the strength the evidence allows, is a reporting norm rather than a universal law: "post-training should be judged not only by its pass@1 but also by how much of the base model's potential it keeps," i.e., coverage metrics like Sharpening Tax should accompany accuracy whenever a lab claims post-training "unlocks" new agentic capability.
Inquiring lines that read this note 13
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems achieve real improvement without external human feedback?- Can constitutional AI training reduce agentic misalignment without task-specific examples?
- Do foundation models develop task-specific shortcuts instead of building stable world models?
- Why do deployed models lack the learning and planning Weinstein attributes to them?
- Do kernel optimization wins show agents discover genuinely novel techniques?
- Does the metaproductivity mismatch occur outside coding-agent benchmarking tasks?
- Does source bias affect real deployed agents or only benchmark environments?
- Why do high-scoring agents default to known techniques rather than novel solutions?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do base models already contain hidden reasoning ability?
Explores whether reasoning capability emerges during pre-training as a latent feature rather than being created by post-training methods like reinforcement learning or fine-tuning.
same elicitation-not-creation claim, extended here from math/reasoning into multi-turn agentic tool use
-
Does RL teach reasoning or just when to use it?
Does reinforcement learning in thinking models actually create new reasoning abilities, or does it simply teach existing capabilities when to activate? This matters for understanding where reasoning truly emerges.
compatible framing: post-training schedules deployment of existing capability rather than installing new capability
-
Can reinforcement learning discover reasoning strategies base models cannot?
Does RL training truly expand what models can do, or does it just find solutions already hidden in base models? ProRL tests this by running RL longer and on diverse tasks beyond mathematics.
tension: prolonged RL with entropy control reportedly beats base models at all pass@k, unlike the standard post-training checkpoints studied here
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
this paper's bimodalization (success rates polarizing to 0 or 1) is a behavioral signature of that same entropy collapse
-
Does the choice of RL algorithm actually matter for reasoning?
Expert Iteration, PPO, and RC-RL show similar performance on reasoning tasks. The question is whether algorithm choice drives results or whether something deeper—like the pretrained model itself—sets the real limits.
evidence for — B shows SFT tightens the pretrained exploration ceiling, explaining why post-trained models lose coverage in A
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Sharpening Tax in Post-Training
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Large Language Models Think Too Fast To Explore Effectively
Original note title
post-training pays a sharpening tax on agentic tasks — base models with a harness outscale post-trained models in solution coverage