Line of inquiry
Inquiring lines›How should we train models for cap…›How do attention and architecture…›this line of inquiry
Can alternative training methods improve on supervised fine-tuning for language models?
A broader line of inquiry — a family of 34 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 34
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can preference learning fix the rigid output format problem better than supervised training?
- How does preference-based training compare to supervised fine-tuning for function calling?
- Can preference model training be redesigned to prioritize factual correction over user agreement?
- Can structured natural language feedback outperform scalar rewards in RL?
- How do relational reward signals compare to absolute preference encodings in RL?
- How do loss functions simultaneously shape both learning and decision quality?
- Can distillation and reward optimization happen in a single training loop?
- Can compact reward function representations beat text based personalization approaches?
- Can light human signals steer already-learned behavior without preference labels?
- How do inference-time reward methods compare to per-user fine-tuning?
- Can smaller judge models better capture human preferences than larger prompted models?
- How do Q-value models improve action selection compared to value models?
- How do binary comparisons constrain reward scale in multi-user preference learning?
- Can reward-guided decoding replace weight fine-tuning for personalized alignment?
- Can input-only training encode user preferences without task-specific labels?
- Can counterfactual data augmentation fully eliminate preference model miscalibration?
- What alignment properties emerge when the reward model disappears?
- Can we reverse the instruction-following deficit through targeted training?
- Can trajectory quality filtering improve model training in noisy environments?
- How do self-generated preference pairs from a strong teacher compare to human feedback?
- How does SDPO relate to agents learning from verbal reflection without parameter updates?
- Can distillation methods extract directional guidance that scalar RL cannot access?
- Can continuous spectrum training outperform sequential SFT-then-RL stages?
- How do pairwise comparisons convert subjective quality into trainable ranking signals?
- Can rich environment feedback replace human preference labels entirely?
- Can negative feedback through critiques achieve the same steering flexibility as positive preferences?
- Can importance sampling reduce variance in off-policy reward estimation?
- How do text-based preference summaries compare to embedding vectors for conditioning?
- Can curriculum degradation of document quality accelerate policy learning?
- What preference dimensions do base reward functions typically capture?
- How does upstream value embedding differ from downstream alignment patches?
- Can information-gain principles improve how we choose what to label?
- How do neural networks extend contextual bandits beyond linear reward assumptions?
- Why does DPO outperform SFT specifically for function calling tasks?