Line of inquiry
Inquiring lines›Where does language-model reasonin…›When and why does chain-of-thought…›this line of inquiry
Do corrupted reasoning traces serve as effective supervision signals?
A broader line of inquiry — a family of 23 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 23
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do corrupted reasoning traces sometimes generalize better than correct ones?
- Why do corrupted traces maintain performance as well as correct traces?
- Why do deliberately corrupted reasoning traces sometimes generalize better than correct ones?
- Do corrupted reasoning traces teach something different than pure success traces?
- Can corrupted reasoning traces be reliably distinguished from correct ones?
- What makes some reasoning traces better supervision than others despite equal accuracy?
- Why do invalid prompts produce reasoning traces as effectively as valid ones?
- Can deliberate corruption of reasoning traces harm out of distribution generalization?
- Why do invalid reasoning steps produce nearly the same performance gains?
- How does post-training on traces improve performance without semantic reasoning?
- What makes reasoning traces effective or ineffective for solving problems?
- Can training on reasoning traces teach actual self-correction or only confident first answers?
- How much do compressed reasoning traces transfer across different models?
- Why do reasoning models produce unfaithful or unhelpful reasoning traces?
- Does reasoning training create blind spots in premise detection?
- How does confidence filtering improve selection of reasoning traces?
- Does reasoning trace style explain why RL post-training improves model reasoning?
- Why does mixing reasoning traces from different teachers destabilize learning?
- Why do wrong numbers cost less accuracy than shuffled reasoning steps?
- Why do human-curated thought examples fail to improve model thinking?
- How does an aggregator use diverse complementary traces to improve final answers?
- What makes a set of traces collectively useful beyond their individual quality?
- Can removing failed branches from edited traces improve previous mistakes?