SYNTHESIS NOTE
TopicsConversation Topics Dialogthis note

Can LLMs truly update shared conversational common ground?

Explores whether large language models can participate symmetrically in Stalnaker's picture of communication, where speakers mutually revise shared assumptions. The question matters because it reveals whether human-LLM dialogue is genuinely interactive or structurally asymmetrical.

Synthesis note · 2026-05-01 · sourced from Conversation Topics Dialog
Why do AI conversations reliably break down after multiple turns? Where exactly do language models fail at structural language tasks?

On Stalnaker's picture, communication is a process of mutually proposing and accepting updates to shared assumptions. Each assertion is a candidate for incorporation into common ground; participants accept, query, or reject. The common ground evolves as conversation proceeds, and that evolution is itself the substance of communication.

LLMs cannot participate in this process symmetrically. The prompt establishes the model's working context, and the model interprets subsequent turns within that frame. Even when a user pivots — shifting from climate policy to historical precedent, or revealing they are not actually a five-year-old after asking for a five-year-old explanation — the LLM cannot smoothly absorb the revision into a jointly held common ground. It either ignores the pivot, fabricates continuity, or requires the user to re-scaffold from scratch. The asymmetry is structural: humans propose, the LLM either adopts or routes around, but the LLM cannot itself propose updates that change what counts as background.

This is a deeper deficit than failures of memory or inference. It means that the conversational scoreboard — Lewis's mechanism for tracking what counts as a felicitous next move — is one-sidedly maintained by the user. The user is keeping score for both players. The model is producing moves that look responsive but cannot reciprocally update the score in the way the conversational practice requires. What looks like dialogue is structurally closer to oracle-consultation, where the questioner provides all context and the oracle returns a response framed within it.

Inquiring lines that read this note 105

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How should dialogue systems represent uncertainty from noisy speech input? Does conversational format create illusions of genuine AI communication? How can language models sustain linguistic synchrony and intersubjectivity during dialogue? How does rhetorical adaptation affect LLM persuasion and detectability? How can LLM user simulators model realistic goal-driven conversation? How do language models establish social grounding in human dialogue? How should dialogue recommender systems manage conversation history and state? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Why should disagreement be treated as signal in collaborative reasoning? How do self-generated feedback mechanisms enable effective model learning? Why do language models struggle with implicit discourse relations? Why do LLM chatbots fail as independent therapeutic agents? What makes dialogue-based explanation more successful than monologue? Can LLM personas constitute genuine psychology or remain linguistic role-play? Does RLHF training sacrifice accuracy and grounding for user agreement? What articulatory information do speech signals carry that text cannot? What distinguishes dynamic from static grounding in dialogue systems? Is embodied interaction necessary for language meaning and genuine agency? How do formal dialogue structures reveal conversation coherence mechanisms? How should conversational agents balance goal-driven initiative with user control? Why do language models reinforce false assumptions instead of correcting them? How should we design LLM systems to maintain alignment and control? What coordination failures limit multi-agent LLM systems as they scale? Do language models learn genuine linguistic structure or just surface patterns? Why do multi-turn conversations degrade AI intent and coherence? How can AI alignment serve diverse human preferences at scale? How do multi-agent systems achieve genuine cooperation and reasoning? How do language models inherit human biases from training data?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Common ground in human-LLM conversation cannot be jointly updated because the LLM treats prompts as static frames