Line of inquiry
Inquiring lines›What determines the reliability an…›How do systems improve effectively…›this line of inquiry
Why does single-turn training fail to generalize to multi-turn tasks?
A broader line of inquiry — a family of 18 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 18
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does single-turn optimization undermine multi-turn collaborative dynamics?
- Why do single-turn RL methods fail to generalize to multi-turn tasks?
- Does longer interaction horizon require fundamentally different evaluation approaches?
- How do complete multi-turn trajectories differ from isolated task examples?
- How does single-turn training undermine multi-turn strategic dialogue?
- How can interactive evaluation avoid replicating fragmentation problems from response-centered benchmark culture?
- How does bounded committed state prevent multi-turn agent failures better than transcript replay?
- Why do next-turn reward objectives fail to encourage multi-turn goal progress?
- What specific metrics distinguish single-turn versus multi-turn collaboration success?
- How does structured environment state compare to transcript replay for multi-turn reasoning?
- What makes session-aware multi-turn tracking necessary for asynchronous training?
- What preference optimization strategy works best for multi-turn social alignment?
- How much does forcing single-choice answers damage alignment with complex intent?
- How does evaluating interaction trajectories change what we measure beyond correctness?
- Why do one-shot transparency studies miss the temporal reversal entirely?
- Can unified policies handle negative feedback and critique transformation simultaneously?
- What path-dependencies lock in AI's societal impacts before they become visible?
- What does tight coupling mean in normal accident theory for AI?