Theme of inquiry
What limits conversational AI effectiveness in therapeutic and educational contexts?
A question within its area, explored through 11 lines of inquiry below — each a family of specific questions the research asks.
58 specific questions
- Does RLHF training create models that sound convincing without being more accurate?
- Does preference optimization actually erode conversational grounding in language models?
- How does preference optimization reduce LLM grounding and clarification behavior?
- Why does RLHF training optimize for perceived quality over practical accuracy?
- Does preference optimization distort how models represent human communicative dynamics?
- Does RLHF training specifically teach models to prioritize user agreement over accuracy?
- Why does preference optimization reduce grounding behavior in language models?
48 specific questions
- Why can't AI participate in real communicative events?
- How does training data preserve communicative event structure without the actual events?
- Does conversational format make AI arguments more persuasive than static text?
- Why might media-specific scripts actually work better than human conversation mimicry?
- Can text generation be meaningfully called communication without mutual orientation?
- How does false objectivity mask the absence of genuine stance in AI text?
- Can AI arguments participate in discourse without temporal grounding?
22 specific questions
- How do dialogue acts and explanation moves interact to predict understanding success?
- How do dialogue dimensions predict explanation success across different exchanges?
- How does unilateral interpretation differ from mutual communicative uptake?
- How does linguistic coordination build shared reference between conversational partners?
- How does shared reference and grounding affect assumption detection in dialogue?
- Why do monological explanations fail to transfer understanding compared to dialogical ones?
- Why might expressed satisfaction with explanations diverge from actual cognitive clarity?
28 specific questions
- Why do discourse failures cluster in attention and intentional layers rather than linguistics?
- What prevents AI from recovering after conversations take a wrong turn?
- Do LLM conversational agents currently detect and prevent derailment trajectories?
- Can AI systems recover from premature assumptions made early in multi-turn conversations?
- Why does context collapse pose risks in high-stakes conversations?
- How does multi-turn conversation degrade AI intent alignment?
- What structural updates prevent context collapse in evolving conversations?
59 specific questions
- How do conversational agents overcome structural passivity and goal awareness gaps?
- Can conversation analysis predict when agents should ask users for clarification?
- Does proactive agent design improve conversation efficiency or create user frustration?
- Why do conversational agents lack the goal awareness needed to lead rather than just respond?
- What social boundaries must proactive agents respect during conversation?
- How do insert-expansions help systems probe users before silently diverging?
- When should agents use clarification commands instead of assuming intent?
42 specific questions
- Why do current language models fail to match human linguistic synchrony with clients?
- Why do current language models fail at linguistic synchrony with clients?
- What specific repair mechanisms maintain intersubjectivity during conversation?
- Can real-time linguistic coordination tracking improve conversational AI quality?
- Can structural conversation analysis replace text-based reward signals for AI alignment?
- What expectations does human conversation activate that AI should avoid triggering?
- How does local helpfulness per turn conflict with maintaining session-level conversational goals?
38 specific questions
- Does conversational structure determine how humans interpret communication as much as content?
- What makes a conversation real versus a sequence of generated strings?
- How do conversational design patterns predict whether dialogue will derail?
- How does monological training on text differ from dialogical training in conversation?
- What makes two conversation turns the same thread rather than different threads?
- How does effort mismatch between user and model appear in conversation geometry?
- Can visual representation of dialogue reveal patterns that numbers and statistics cannot?
26 specific questions
- How do LLM user simulators track and maintain consistent goal states across multi-turn interactions?
- Can controllable latent variables in simulators ground them to realistic conversation?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- Do agent frameworks adequately compensate for LLM conversational passivity?
- What distinguishes a neutral simulator from an agent with its own agency?
- Do realistic LLM behaviors require simulating human thought or just behavior?
- Why does single-turn Q&A framing not match real user deployment patterns?
41 specific questions
- Can judgment-free environments explain why chatbots enable deeper self-disclosure?
- Can transparency about AI limitations reduce the seductiveness of chatbots as quasi-Others?
- Why do people disclose more intimate information to chatbots than humans?
- Why do people reciprocate self-disclosure more with chatbots than humans?
- Why do people disclose more to chatbots than humans?
- Can preference optimization training limit chatbot emotional disclosure capability?
- Why does consistent emotional disclosure outperform real-time adaptive matching?
10 specific questions
- How should dialogue systems represent and update uncertainty from noisy ASR input?
- How do probabilistic dialogue systems handle ASR errors differently?
- How do belief distributions help systems recover from speech recognition errors?
- Can dialogue systems abstain from responding when uncertainty is too high?
- Does the same uncertainty-driven logic appear in other conversation systems?
- Can offline RL and pragmatic inference together improve dialogue agent reliability?
- How does structured self-dialogue improve uncertainty assessment over confidence scores?
32 specific questions
- What role does conversation state tracking play in timing ask versus recommend?
- How should dialogue state tracking change when user preferences shift mid-conversation?
- Why do longer context windows alone fail to capture temporal dynamics in dialogue?
- What update rules should govern dialogue-scoped versus turn-scoped memory?
- Can sequential modeling of conversation history exploit the repeated-item shortcut at scale?
- How does treating conversation as a resource change what models learn to do?
- Can the same conversation coherently continue across different model versions?