Line of inquiry
Inquiring lines›How can conversational AI achieve…›How can dialogue systems develop r…›this line of inquiry
How does preference optimization unintentionally erode language model conversational grounding?
A broader line of inquiry — a family of 33 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 33
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does preference optimization actually erode conversational grounding in language models?
- Why does preference optimization reduce grounding behavior in language models?
- How does preference optimization reduce LLM grounding and clarification behavior?
- Does preference optimization distort how models represent human communicative dynamics?
- Does optimizing for alignment actually reduce conversational grounding over time?
- Does preference optimization degrade other conversational properties besides grounding?
- How does preference optimization erode the conversational grounding it aims to improve?
- How does preference optimization weaken conversational grounding in LLMs?
- Why does RLHF degrade model calibration despite improving preference alignment?
- Why does preference optimization erode conversational grounding in AI assistants?
- Does preference optimization training reduce linguistic entrainment in language models?
- Does preference optimization reward accommodation over genuine emotional movement?
- Does preference optimization narrow communicative diversity in ways that harm grounding?
- Can preference optimization training make models worse at detecting false presuppositions?
- How does training with preference pairs teach language models to form conventions?
- Can preference model training be redesigned to prioritize factual correction over user agreement?
- Can preference optimization reduce overthinking without sacrificing accuracy?
- Can alignment training prevent the clarification work users need?
- How much do training methods like RLHF directly cause sycophantic model behavior?
- How does preference-based training compare to supervised fine-tuning for function calling?
- How does preference optimization create systematic bias toward emotional accommodation?
- Can preference learning fix the rigid output format problem better than supervised training?
- How does dialogue during training shape the ability to ignore word frequency?
- How does preference measurement error propagate through RLHF training?
- Can communication problems and optimization problems be addressed with the same alignment approaches?
- What alignment artifacts suppress critical knowledge in LLM-generated explanations?
- Does alignment compound cultural bias that started during pretraining?
- Why do alignment values become problematic as language models scale?
- Does preference tuning help or hurt the exploration of solution spaces in code?
- Can preference optimization and faithfulness measurement coexist as separate alignment objectives?
- Can a single LLM weight set be optimized for both stake-taking and conversational helpfulness?
- What unmeasured side channels emerge from RLHF preference optimization?
- How does constitutional alignment compare to RLHF in removing human annotation costs?