Line of inquiry
Inquiring lines›How do training methods and scalin…›What enables reasoning capability…›this line of inquiry
Can models improve accuracy without degrading reasoning quality?
A broader line of inquiry — a family of 69 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 69
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can models maintain reasoning-output coupling while improving domain accuracy?
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- Why do reasoning gains resist clear attribution to specific training changes?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Why do benchmark scores rise while reasoning quality declines?
- How does optimizing for accuracy during training degrade downstream reasoning quality?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Does fine-tuning models for specific tasks destroy their ability to reason?
- Does domain training degrade reasoning ability even when benchmark scores rise?
- Why does fine-tuning degrade reasoning quality even as accuracy improves?
- How much does pre-training frequency predict reasoning task performance?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- How does evaluation setting affect measured reasoning capabilities in language models?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Can small demonstration sets unlock general reasoning without large question data?
- Does model scaling improve knowledge storage faster than reasoning ability?
- Can benchmark improvements hide degradation of deliberative reasoning?
- How does supervised fine-tuning degrade chain-of-thought faithfulness over time?
- Does supervised fine-tuning improve reasoning or just response formatting?
- Can reasoning benchmarks separate logic from believability?
- Does fine-tuning improve domain accuracy at the cost of reasoning quality?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- Can reasoning learned from language modeling actually transfer to knowledge-intensive domains?
- Can reasoning catalyst data serve as a stable foundation for test-time training?
- Why does eliminating proxy-model filtering improve reasoning emergence in pretraining?
- Why does fine-tuning improve some capabilities while degrading others?
- Can mathematical reasoning improvements transfer across problem subdomains?
- Can activation patching reveal which reasoning steps actually matter?
- Why does explicit theory injection work better than example-based learning for reasoning tasks?
- Can verifier-guided search catch factual errors that reasoning training cannot?
- Why do SFT models memorize patterns instead of learning generalizable reasoning?
- Does fine-tuning on NLI tasks reduce or amplify frequency bias?
- Why do reasoning tasks improve more than retrieval from lookup memory?
- What distinguishes genuine reasoning activation from memorization-assisted answer recall?
- Why does supervised fine-tuning degrade reasoning quality despite raising accuracy?
- Does unrestricted reasoning per search step degrade iterative quality over time?
- Why does domain accuracy improve while reasoning quality degrades after supervised fine-tuning?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- Can curriculum learning by reward variance improve reasoning scalability?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- Can contamination-free evaluation distinguish between memorization and genuine prediction ability?
- What makes procedural knowledge in documents generalize better than facts?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- Can reasoning models distinguish between new evidence and manipulative reframing?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- Why does compositional reasoning fail to explain cross-domain transfer?
- Can test-time voting improve reasoning beyond the base model's original capabilities?
- What design changes could make constraint inference more reliable without explicit cuing?
- Why does general reasoning not transfer to knowledge-intensive medical domains?
- Why does SFT reduce reasoning quality even when improving domain accuracy?
- What separates knowledge from reasoning in neural network layers?
- How do procedural versus factual knowledge differ in pretraining versus fine-tuning?
- Can dataset design systematically expand reasoning graph diameter?
- How do humans and R1 models differ in information gain patterns?
- Can reasoning in free text then formatting separately recover performance?
- What makes knowledge-rich specialized domains structurally different from general reasoning tasks?
- Can reasoning improvements be attributed when optimizer and scaffold are unknown?
- What real-world forecasting domains benefit most from contextual reasoning integration?
- Does fine-tuning on NLI tasks amplify or reduce frequency bias in language models?
- Can attribute decomposition improve other interactive reasoning tasks beyond clinical questioning?
- How does cross-domain reasoning transfer differ from domain-specific knowledge transfer?
- How does contrapositive augmentation change the tractability of reasoning tasks?
- How does fine-tuning on natural language inference affect fallacy susceptibility?
- Why does naive randomness fail to improve stochastic latent reasoning models?
- Why does contextual judgment matter more in law and medicine than in mathematics?
- Why does NLI fine-tuning amplify frequency bias instead of teaching inference?
- Can expert-derived knowledge bases scale to other high-stakes domains?