Line of inquiry
Inquiring lines›How should we train models for cap…›How do different training strategi…›this line of inquiry
What pretraining choices and baseline capability constrain reinforcement learning gains?
A broader line of inquiry — a family of 49 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 49
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can single-problem fine-tuning match full RL pipeline reasoning gains?
- Does format-based pretraining determine how models respond to reinforcement learning?
- Why does RL improve sampling efficiency but not expand capability boundaries?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- What makes pretraining composition more important than reward engineering?
- How do sparse parameter updates enable when-not-how training to work?
- How does baseline capability level affect RL improvement ceiling?
- Can out-of-distribution tests expose memorization in reinforcement learning fine-tuned models?
- Can smaller models achieve domain expertise through focused RL training?
- How does post-training shift models from passive prediction to on-policy action?
- What role does natural language play in breaking reinforcement learning performance plateaus?
- Does the pretrained model prior limit RL search capability more than the optimization algorithm itself?
- Can in-context learning replicate the timing effects that RL teaches models?
- How does imitation pretraining followed by RL exploration compare to either method alone?
- Why does the pretrained prior determine the exploration ceiling?
- What scaling properties emerge from RL training dynamics beyond verification?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- Can in-context reinforcement learning match human sample efficiency on real problems?
- How does pretraining determine what RL can later teach a model?
- Can multi-turn reinforcement learning improve tool use in language models?
- What capacity threshold determines whether RL teaches activation versus shortcut learning?
- Why does multi-turn RL generate orders of magnitude more tokens than single-turn?
- Can RL directly optimize attention distributions instead of text generation?
- What breaks when you apply reinforcement learning after supervised fine-tuning?
- How does LLM simulation of APIs avoid instability without sacrificing training signal?
- Can LLM-synthesized behavioral heuristics compete with learned policy improvements?
- Why do models follow a two-phase pattern of procedural then strategic learning?
- Does sparsity in RL arise from training on policy-distribution data?
- How should humans specify deterministic abstractions of RL problems?
- Why do single-turn RL methods fail to generalize to multi-turn tasks?
- What training duration is actually needed for RL to expand capabilities?
- Can offline reinforcement learning improve dialogue policy baseline performance?
- Why does reinforcement learning training degrade model calibration?
- Can the exploration ceiling be raised beyond what pretraining established?
- How does trajectory filtering handle noise when language models use code execution tools?
- Which recipe choices determine the asymptotic ceiling in RL training?
- How do residual connections and layer norm stabilize training in deep RL?
- Does RL training redirect self-doubt into productive gap analysis?
- How does absolute-advantage weighting concentrate training on boundary cases?
- Why does early experience provide better warm-starts for downstream reinforcement learning?
- How do RL training and base models differ in creating MI peaks?
- What makes software engineering environments better suited for RL than other interactive domains?
- What makes a task at the edge of competence optimal for RL?
- Do disorder-specific RL policies outperform single policies across anxiety, depression, and schizophrenia?
- Why does online RL succeed where supervised training fails for self-correction?
- Can meta-reinforcement learning explain why this bias pattern emerges rationally?
- What makes session-aware multi-turn tracking necessary for asynchronous training?
- How do RL subnetworks identified from different random seeds compare?
- How does behavior cloning reduce complexity before RL training in rerankers?