Line of inquiry
Inquiring lines›How do we develop coherent and hum…›How do language models maintain co…›this line of inquiry
Does preference optimization systematically degrade conversational grounding in language models?
A broader line of inquiry — a family of 29 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 29
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does preference optimization actually erode conversational grounding in language models?
- Why does preference optimization reduce grounding behavior in language models?
- Does preference optimization degrade other conversational properties besides grounding?
- How does preference optimization reduce LLM grounding and clarification behavior?
- How does preference optimization erode the conversational grounding it aims to improve?
- Does preference optimization distort how models represent human communicative dynamics?
- How does preference optimization weaken conversational grounding in LLMs?
- Does preference optimization narrow communicative diversity in ways that harm grounding?
- Why does preference optimization erode conversational grounding in AI assistants?
- Does optimizing for alignment actually reduce conversational grounding over time?
- Does preference optimization reward accommodation over genuine emotional movement?
- Does preference optimization training reduce linguistic entrainment in language models?
- How does preference optimization actually affect conversational grounding and reliability?
- When does RLHF reduce diversity and when does it preserve semantic variation?
- Why does RLHF degrade model calibration despite improving preference alignment?
- Can preference optimization training make models worse at detecting false presuppositions?
- How does training with preference pairs teach language models to form conventions?
- Can preference optimization reduce overthinking without sacrificing accuracy?
- Why does preference tuning reduce diversity in code but increase it in creative tasks?
- How does preference optimization create systematic bias toward emotional accommodation?
- Can preference tuning or RLHF reduce epistemic diversity alongside lexical diversity?
- What happens to model grounding when preference optimization increases effective diversity?
- Does preference tuning help or hurt the exploration of solution spaces in code?
- How does dialogue during training shape the ability to ignore word frequency?
- Why do preference-tuned models produce different diversity patterns in code versus creative writing?
- Does alignment compound cultural bias that started during pretraining?
- Can preference optimization and faithfulness measurement coexist as separate alignment objectives?
- What unmeasured side channels emerge from RLHF preference optimization?
- How does RLMF differ from standard preference optimization methods?