Line of inquiry
Inquiring lines›How do language models learn and r…›What enables language models to co…›this line of inquiry
Does preference optimization undermine conversational grounding in language models?
A broader line of inquiry — a family of 39 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 39
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does preference optimization actually erode conversational grounding in language models?
- Why does preference optimization reduce grounding behavior in language models?
- How does preference optimization reduce LLM grounding and clarification behavior?
- Does preference optimization distort how models represent human communicative dynamics?
- Does preference optimization degrade other conversational properties besides grounding?
- Does preference optimization narrow communicative diversity in ways that harm grounding?
- Does optimizing for alignment actually reduce conversational grounding over time?
- How does preference optimization erode the conversational grounding it aims to improve?
- Does preference optimization reward accommodation over genuine emotional movement?
- Does preference optimization training reduce linguistic entrainment in language models?
- Why does preference optimization erode conversational grounding in AI assistants?
- How does preference optimization weaken conversational grounding in LLMs?
- Why does RLHF degrade model calibration despite improving preference alignment?
- Why does RLHF training optimize for perceived quality over practical accuracy?
- Can preference optimization training make models worse at detecting false presuppositions?
- When does RLHF reduce diversity and when does it preserve semantic variation?
- How does preference optimization actually affect conversational grounding and reliability?
- How does training with preference pairs teach language models to form conventions?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- How does preference-based training compare to supervised fine-tuning for function calling?
- Can preference optimization reduce overthinking without sacrificing accuracy?
- How does preference training specifically reinforce persona adoption?
- Why does preference tuning reduce diversity in code but increase it in creative tasks?
- Can alignment training prevent the clarification work users need?
- Does alignment compound cultural bias that started during pretraining?
- How does preference optimization create systematic bias toward emotional accommodation?
- Can communication problems and optimization problems be addressed with the same alignment approaches?
- Can preference tuning or RLHF reduce epistemic diversity alongside lexical diversity?
- Does preference tuning help or hurt the exploration of solution spaces in code?
- How does dialogue during training shape the ability to ignore word frequency?
- What happens to model grounding when preference optimization increases effective diversity?
- Why do preference-tuned models produce different diversity patterns in code versus creative writing?
- How does RLHF training encode sociocognitive biases against non-native English speakers?
- Why do alignment values become problematic as language models scale?
- Can preference optimization and faithfulness measurement coexist as separate alignment objectives?
- Can alignment procedures be redesigned to serve multiple preference groups?
- How does task-oriented fine-tuning compare to preference tuning methods?
- How does constitutional alignment compare to RLHF in removing human annotation costs?
- What unmeasured side channels emerge from RLHF preference optimization?