Line of inquiry
Inquiring lines›How can we optimize language model…›How do test-time resources and tra…›this line of inquiry
Does fine-tuning sacrifice reasoning ability to improve task accuracy?
A broader line of inquiry — a family of 38 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 38
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does fine-tuning degrade reasoning quality even as accuracy improves?
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- Does SFT degrade reasoning quality while improving domain accuracy?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Why does supervised fine-tuning degrade reasoning quality despite raising accuracy?
- Does supervised fine-tuning improve reasoning or just response formatting?
- Does fine-tuning models for specific tasks destroy their ability to reason?
- How does supervised fine-tuning degrade chain-of-thought faithfulness over time?
- Does reasoning fine-tuning actually damage a model's ability to abstain?
- Does reasoning fine-tuning actually reduce a model's ability to abstain?
- Does fine-tuning push models toward reasoning shortcuts that bypass the chain entirely?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- Why does reasoning fine-tuning reduce a model's ability to abstain?
- Does fine-tuning improve domain accuracy at the cost of reasoning quality?
- Why does domain accuracy improve while reasoning quality degrades after supervised fine-tuning?
- Why does fine-tuning improve some capabilities while degrading others?
- Can reasoning fine-tuning improve both capability and instruction compliance together?
- How does optimizing for accuracy during training degrade downstream reasoning quality?
- Why does SFT reduce reasoning quality even when improving domain accuracy?
- Why do SFT models memorize patterns instead of learning generalizable reasoning?
- Why does reasoning fine-tuning reduce model abstention capacity by 24 percent?
- Why does fine-tuning sometimes damage chain-of-thought reasoning even when accuracy improves?
- Does domain training degrade reasoning ability even when benchmark scores rise?
- Why does reasoning fine-tuning suppress the confidence signals that adaptive retrieval needs?
- Why does fine-tuning models for continuous reasoning cause catastrophic forgetting?
- How does preference learning differ from supervised finetuning for reasoning?
- Do models trained for safety over-refuse compared to models trained for reasoning?
- Does fine-tuning on NLI tasks reduce or amplify frequency bias?
- Does fine-tuning on NLI tasks amplify or reduce frequency bias in language models?
- Why does full multi-task fine-tuning perform worse than sequential training?
- How does fine-tuning on natural language inference affect fallacy susceptibility?
- Why does SFT fail when expert demonstrations are too long for small models?
- What hidden costs emerge when you fine-tune models for a single domain?
- Why does fine-tuning fail to remove temporal contamination from pretraining?
- Can reasoning improvements be attributed when optimizer and scaffold are unknown?
- How does task-oriented fine-tuning compare to preference tuning methods?
- Why does NLI fine-tuning amplify frequency bias instead of teaching inference?