Line of inquiry
Inquiring lines›How can we optimize language model…›How do test-time resources and tra…›this line of inquiry
Why doesn't reasoning ability transfer from general to specialized domains?
A broader line of inquiry — a family of 27 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 27
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can reasoning learned from language modeling actually transfer to knowledge-intensive domains?
- Does model scaling improve knowledge storage faster than reasoning ability?
- Why does reasoning training improve math but hurt knowledge tasks?
- Can mathematical reasoning improvements transfer across problem subdomains?
- What makes knowledge-rich specialized domains structurally different from general reasoning tasks?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- Why does general reasoning not transfer to knowledge-intensive medical domains?
- What makes procedural knowledge in documents generalize better than facts?
- What is the difference between procedural knowledge and factual retrieval in reasoning?
- Can we transfer reasoning structure without copying surface form?
- Why do reasoning tasks improve more than retrieval from lookup memory?
- How do procedural versus factual knowledge differ in pretraining versus fine-tuning?
- How does cross-domain reasoning transfer differ from domain-specific knowledge transfer?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- What separates knowledge from reasoning in neural network layers?
- Why does contextual judgment matter more in law and medicine than in mathematics?
- Why does reasoning transfer across different numbers but factual recall does not?
- How many document exposures does procedural knowledge versus factual information require?
- How does the knowing-doing gap widen as tasks become more complex?
- Why do medical and mathematical tasks require fundamentally different model capabilities?
- What does pass@k reveal about base model reasoning capacity?
- Can reasoning skills trained on law improve performance in STEM?
- Can expert-derived knowledge bases scale to other high-stakes domains?
- Do reasoning languages like Prolog follow the same two-constraint transfer pattern?
- What makes symbolic operations different from general knowledge questions?
- How much of MATH-500 improvement comes from data contamination versus real reasoning gains?
- Why do readability and style metrics plateau while reasoning improves with scale?