Line of inquiry
Inquiring lines›What drives capability improvement…›What training and inference approa…›this line of inquiry
How does fine-tuning trade off accuracy against reasoning quality?
A broader line of inquiry — a family of 82 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 82
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can models maintain reasoning-output coupling while improving domain accuracy?
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Does domain training degrade reasoning ability even when benchmark scores rise?
- Does fine-tuning models for specific tasks destroy their ability to reason?
- How does optimizing for accuracy during training degrade downstream reasoning quality?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Why does fine-tuning improve some capabilities while degrading others?
- Why does fine-tuning degrade reasoning quality even as accuracy improves?
- Do instruction-tuned models learn tasks or just output format distributions?
- Can models recover knowledge with completely unrelated retraining tasks?
- Does fine-tuning improve domain accuracy at the cost of reasoning quality?
- Can in-context learning substitute for domain-specific training altogether?
- How does supervised fine-tuning degrade chain-of-thought faithfulness over time?
- Why does explicit theory injection work better than example-based learning for reasoning tasks?
- Can reasoning learned from language modeling actually transfer to knowledge-intensive domains?
- Does specialized training in one domain create capability cliffs elsewhere?
- How does evaluation setting affect measured reasoning capabilities in language models?
- Why does eliminating proxy-model filtering improve reasoning emergence in pretraining?
- Why do language models substitute parametric knowledge over retrieved context mid-reasoning?
- Why do SFT models memorize patterns instead of learning generalizable reasoning?
- Why does compositional reasoning fail to explain cross-domain transfer?
- Can mathematical reasoning improvements transfer across problem subdomains?
- What makes procedural knowledge in documents generalize better than facts?
- Does supervised fine-tuning improve reasoning or just response formatting?
- Why does domain accuracy improve while reasoning quality degrades after supervised fine-tuning?
- Does finetuning facts into weights overwrite existing model capabilities?
- Does optimizing directly for semantic diversity improve both reasoning quality and exploration?
- Can reasoning catalyst data serve as a stable foundation for test-time training?
- Why do foundation models develop heuristics instead of world models?
- Can training on diverse related tasks be more efficient than task-specific training?
- Why do explicit quality criteria outperform learning quality from examples alone?
- Can instance seeds work for tasks beyond language understanding benchmarks?
- What makes some contexts learnable as rules versus requiring model retraining?
- Does fine-tuning on NLI tasks reduce or amplify frequency bias?
- Why does capturing domain structure reduce data requirements more than raw volume?
- Why does supervised fine-tuning degrade reasoning quality despite raising accuracy?
- What training cost tradeoffs exist between fine-tuning and other knowledge injection methods?
- Why does structuring knowledge into taxonomies outperform larger unorganized training sets?
- Where does skill extraction fail compared to genuine model adaptation?
- Why does SFT reduce reasoning quality even when improving domain accuracy?
- Do base models or search agents win at open-ended prediction tasks?
- Why does semantic similarity retrieval enable skill transfer to novel situations?
- How do humans and R1 models differ in information gain patterns?
- Does knowledge structure matter more than knowledge volume for model training?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- How does cross-domain reasoning transfer differ from domain-specific knowledge transfer?
- What real-world forecasting domains benefit most from contextual reasoning integration?
- Does training on granular tasks beat training on the full function calling problem?
- What makes knowledge-rich specialized domains structurally different from general reasoning tasks?
- How do foundation models develop task-specific heuristics instead of world models?
- Which medical tasks benefit most from domain-adaptive pretraining?
- Why does fine-tuning change how models process retrieved context?
- Do text-space skills transfer learning across different frontier models?
- Does fine-tuning on NLI tasks amplify or reduce frequency bias in language models?
- Why do non-experts default to familiar chart types despite domain complexity?
- How do training-time and inference-time knowledge injection techniques compare?
- Why does fine-tuning fail to remove temporal contamination from pretraining?
- Do newer language model generations improve forecasting ability without additional training?
- Can dense models partially address modality friction without full expert specialization?
- How does retrieval-augmented training reduce domain specialization cliff failures?
- Which domains need knowledge injection versus reasoning-focused training?
- Why does naive randomness fail to improve stochastic latent reasoning models?
- Can learned priors effectively select and weight ensemble members by inference budget?
- Can attribute decomposition improve other interactive reasoning tasks beyond clinical questioning?
- How should rapidly evolving domains choose knowledge injection methods?
- How should skill libraries coordinate with gradient-based weight optimization?
- Can the same description-then-retrieve pattern work for domain adaptation without target data?
- Why does contextual judgment matter more in law and medicine than in mathematics?
- Can expert-derived knowledge bases scale to other high-stakes domains?
- What makes hierarchical reasoning effective for taxonomy induction?
- Can extracted skills transfer effectively across different domains and model architectures?
- How robust and general are causal edits like ROME across different facts?
- What knowledge injection routes trade flexibility against training cost?
- What techniques work best for injecting domain knowledge at training time?
- How does cognitive fit theory explain why different tasks need different knowledge structures?
- Why does NLI fine-tuning amplify frequency bias instead of teaching inference?
- How does upward distillation transfer knowledge from smaller to larger networks?
- What makes certain bond distributions more learnable than others?