INQUIRING LINE

Why does handing an AI the correct answer and asking it to explain its way there teach it more than letting it keep guessing?

Why does rationalization work better than just generating more candidate rationales?

This explores STaR's 'rationalization' trick, where a model that got a problem wrong is shown the correct answer and asked to explain how to reach it, and why that helps more than letting it try the problem many more times and keeping whatever happens to come out right.


This explores why giving a model the answer and asking it to explain its way there teaches more than letting it keep guessing. The starting point is STaR Can models improve by filtering only on answer correctness?. The model writes its own step-by-step explanations, only the ones that land on the correct answer are kept, and it is fine-tuned on those. That loop has a blind spot. If the model never solves a problem, it never produces a correct explanation for it, so it never trains on it. The hardest problems, where learning would matter most, drop out of the training data. Rationalization fills that gap by giving the model the answer as a hint and asking it to work backward to a plausible explanation. A problem that taught nothing now produces a training example. One caveat: the corpus summary of STaR covers the filtering-by-correctness core, and the corpus doesn't include a direct head-to-head between rationalization and more sampling. The rest of this answer draws on nearby notes to explain why more sampling is a weak substitute.

The clearest reason comes from work on longer thinking. Does extended thinking actually improve reasoning or just increase variance? finds that longer reasoning traces mostly widen the spread of answers a model produces. They don't make the reasoning itself better. Some answers happen to be right because the spread is wider, and past a certain point the spread gets so wide that accuracy falls. Drawing more candidate explanations works the same way. It catches correct answers the model could already almost produce, and it does nothing for problems outside that reach. Rationalization puts the correct answer into the process from outside, which sampling can't do.

A second reason is that some problems can't be solved by many short independent tries. When does sequential reasoning beat parallel voting? shows that on problems that need to build up intermediate results step by step, one connected chain of reasoning beats voting across many parallel attempts, and the gap grows exponentially. More samples don't add up to one coherent path. Rationalization hands the model the destination, so it only has to build that path toward a known endpoint, a far easier search than wandering toward an unknown one.

That ties into what goes wrong when models explore on their own. Why do reasoning models abandon promising solution paths? finds that reasoning models drift into invalid lines of thought and drop promising ones too early. Good solutions are within reach, but the model leaves them behind. Can abstractions guide exploration better than depth alone? makes a related point: compute spent on a variety of high-level strategies does better than sampling more full solutions in parallel. The common lesson is that guiding the search, whether with an abstraction or with the answer, does more than adding raw attempts. Rationalization is the most direct version of that guidance.

There is a less obvious upshot for anyone curious about self-improving models. Training only on correctness can quietly cap a model at what it can already do, because it only learns from its own successes. Rationalization is one of the earliest and simplest ways to push past that cap: a small hint from outside lets the model learn from its failures as well as its wins.


Sources 5 notes

Can models improve by filtering only on answer correctness?

STaR demonstrates that self-generated rationales filtered exclusively by answer correctness improve reasoning performance significantly. On CommonsenseQA, this correctness-filtered approach achieved 72.5% accuracy, outperforming direct answer fine-tuning and closing the gap with models 30 times larger.

Does extended thinking actually improve reasoning or just increase variance?

Longer thinking traces improve accuracy through variance expansion—broader output distributions cover correct answers more often—not through better reasoning. Beyond a critical threshold, the distribution becomes too diffuse and accuracy drops, revealing the mechanism is sampling coverage, not genuine reasoning improvement.

When does sequential reasoning beat parallel voting?

On structured tasks requiring sequential multi-step reasoning like graph connectivity, chain-of-thought achieves exponentially higher accuracy than parallel voting. The difference emerges because solutions genuinely require accumulating intermediate results sequentially, which short parallel chains cannot achieve.

Why do reasoning models abandon promising solution paths?

Reasoning LLMs exhibit two reinforcing failures: wandering (invalid exploration) and underthinking (premature path-switching). Decoding-level interventions like thought-switching penalties improve accuracy without fine-tuning, suggesting viable solutions exist but are abandoned prematurely.

Can abstractions guide exploration better than depth alone?

RLAD jointly trains abstraction and solution generators, showing that allocating test-time compute to diverse abstractions outperforms parallel solution sampling at large budgets. Abstractions create structured breadth-first exploration that prevents the underthinking failure mode of depth-only reasoning chains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.