Line of inquiry
Inquiring lines›How should agents manage and coord…›How effectively can inference-time…›this line of inquiry
Why does supervised fine-tuning improve accuracy while degrading reasoning quality?
A broader line of inquiry — a family of 25 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 25
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Does supervised fine-tuning improve reasoning or just response formatting?
- Why does fine-tuning degrade reasoning quality even as accuracy improves?
- Why does domain accuracy improve while reasoning quality degrades after supervised fine-tuning?
- Why does SFT reduce reasoning quality even when improving domain accuracy?
- How does supervised fine-tuning degrade chain-of-thought faithfulness over time?
- Why does supervised fine-tuning degrade reasoning quality despite raising accuracy?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Can reinforcement learning fix the reasoning gaps that supervised fine-tuning misses?
- Why do SFT models memorize patterns instead of learning generalizable reasoning?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Does fine-tuning improve domain accuracy at the cost of reasoning quality?
- Why does fine-tuning improve some capabilities while degrading others?
- How does preference learning differ from supervised finetuning for reasoning?
- What makes supervised fine-tuning worsen RL exploration later?
- Why does fine-tuning sometimes damage chain-of-thought reasoning even when accuracy improves?
- Why does SFT fail when expert demonstrations are too long for small models?
- How does non-reasoning SFT prevent overfitting before RL training begins?
- How does task-oriented fine-tuning compare to preference tuning methods?
- How does fine-tuning on natural language inference affect fallacy susceptibility?
- Why does DPO create introspective detection circuits but SFT does not?
- Does fine-tuning on NLI tasks amplify or reduce frequency bias in language models?
- Why does KTO skip supervised fine-tuning while DPO cannot?
- Can demo placement be tuned as a task-specific hyperparameter?