INQUIRING LINE

Bigger AI models can flip the script: untrained versions start beating their fine-tuned selves if you let them try enough times — but does size change when?

How does model scale affect the crossover point between base and post-trained performance?

This explores whether bigger models change the point where a plain base model, given enough attempts, starts solving more problems than its post-trained (instruction- or RL-tuned) version. The corpus describes the crossover itself clearly but has no direct study of how model size moves it.


This explores whether bigger models change the point where a plain base model, given enough attempts, starts solving more problems than its post-trained version. The corpus answers the first half well and the second half only indirectly. The crossover itself is well documented. Across 14 base/post-trained model pairs on agentic benchmarks, base models given only a loose system prompt start out behind. As the number of attempts grows, they eventually solve more distinct tasks than the tuned versions Do base models find more solutions than post-trained ones?. The reason is that post-training splits tasks into two groups: ones the model now solves reliably and ones it never solves. It 'sharpens' toward the easy cases and cuts off rare solutions the base model could still reach occasionally. No note in this collection tracks how that crossover point moves as models get bigger.

The nearby evidence suggests scale changes *what* post-training sharpens toward, not just how much. When RL post-training is run under controlled conditions, it quickly settles on one output format that the model learned during pretraining and suppresses the others. Which format wins depends on model size, and the winner isn't necessarily the best-performing one Does RL training collapse format diversity in pretrained models?. So the alternatives that get lost differ with scale, and the place where the base model's broader range of attempts catches up probably differs too. The collapse also happens in the first training epoch, which suggests the tax arrives early rather than building up slowly.

There's also a reason to expect larger base models to have more to lose. Fine-tuning gains follow a multiplicative scaling law in which a bigger base model helps more than extra pretraining data How should finetuning scale with model and data size?. Emulated fine-tuning experiments separate the two effects: scaling pretraining mostly adds factual knowledge, and scaling fine-tuning mostly adds helpful behavior Do pretraining and fine-tuning scale independently in language models?. Taken together, these suggest a bigger base model holds a larger pool of reachable solutions, and post-training mainly governs how that pool is expressed. That's an inference, not a measured result. But it would mean the cost of sharpening, and the payoff of simply sampling the base model many times, could both grow with size.

Some notes suggest the tax isn't fixed. Training on nearly impossible problems makes it worse: rare lucky successes get rewarded as if they were skill, and the resulting shortcuts damage abilities the model already had Do overly hard RLVR samples actually harm model capabilities?. Methods that keep the tuned model close to the base model's output distribution preserve its ability to learn new tasks later Does staying close to the base model preserve learning ability?. The narrowing also depends on the domain. Preference tuning reduces variety in code but increases it in creative writing Does preference tuning always reduce diversity the same way?. So the crossover point likely depends on the task and the training recipe as much as on parameter count.

The practical upshot: whether the base model wins is a question of attempt budget and tooling, not just of which model is 'better.' A related result shows that wrapping a weaker model in a well-designed harness built by a stronger model nearly doubled its scores without any retraining Can a stronger model lift a weaker one at test time without retraining?. If you want the scale question answered directly, the corpus doesn't have that study yet. It would mean repeating the 14-pair base-vs-post-trained comparison across a ladder of model sizes.


Sources 8 notes

Do base models find more solutions than post-trained ones?

Across 14 model pairs and three agentic benchmarks, base models equipped with only relaxed system prompts eventually surpass post-trained counterparts in pass@K coverage as rollout budget grows. Post-training bimodalizes task outcomes, sharpening performance on easy cases while eliminating rare-but-reachable solutions.

Does RL training collapse format diversity in pretrained models?

Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.

How should finetuning scale with model and data size?

Systematic experiments across 1B–16B models reveal finetuning follows a power-based multiplicative scaling law. Larger base models improve finetuning more than more pretraining data, while increasing PET parameters provides minimal benefit.

Do pretraining and fine-tuning scale independently in language models?

Emulated Fine-Tuning reveals that scaling pretraining improves factual knowledge while scaling fine-tuning improves behavioral helpfulness. This decoupling has architectural roots: pretraining enriches lower-layer knowledge storage, while fine-tuning modifies upper-layer behavior expression.

Do overly hard RLVR samples actually harm model capabilities?

Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.

Show all 8 sources
Does staying close to the base model preserve learning ability?

FST-trained models stay up to 70% closer to their base distribution than parameter-only RL, and this reduced drift preserves the model's ability to learn subsequent tasks effectively. Parameter-only approaches stall when task domains change, while low KL drift enables sustained adaptation.

Does preference tuning always reduce diversity the same way?

RLHF reduces lexical-syntactic diversity in code generation but increases it in creative writing. The direction depends on what each domain incentivizes: code rewards convergence toward correct solutions, while creative writing rewards stylistic distinctiveness.

Can a stronger model lift a weaker one at test time without retraining?

A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.