Can an AI learn tact, saving face, and reading between the lines just from text, without ever really talking to anyone?
Can pragmatic competence emerge from text exposure without interactive grounding?
This explores whether the skill of using language appropriately in context (reading between the lines, saving face, checking you've been understood) can be absorbed from text alone, or whether it needs live back-and-forth with people.
This explores whether the skill of using language appropriately in context can be absorbed from text alone, or whether it needs live back-and-forth with people. The corpus points to a split verdict: text exposure teaches pragmatic habits, but not pragmatic competence. The habits are real. Why do language models avoid correcting false user claims? finds that models often fail to reject a false claim from a user even when they answer the same fact correctly on a direct question. They avoid explicit correction to keep the conversation smooth, which copies a human social norm picked up from training data. Nobody taught that behavior. It emerged from reading. But the model learned the norm without learning when it applies, so it protects the user's feelings at the cost of the truth. Can language models learn meaning without engaging the world? makes the optimistic case: relational structure compressed from text is enough for fluent, culturally situated discourse with no external referents at all.
The habits break down where pragmatics needs context. Can language models adapt implicature to conversational context? tested ChatGPT on scalar implicature, where saying "some" usually implies "not all". The model computed it the same way whether it was told to be literal, whether the information focus shifted, or whether the situation was face-threatening. Humans adjust on all three. Text records what people said after they weighed what was at stake between them, but not the weighing. Why don't language models develop conversation maintenance skills? explains part of why: repairing a reference or handing off a topic is social action, not information transfer. Training rewards predicting information, so nothing pushes those skills to develop. Can language models learn meaning from text patterns alone? states the strongest version. Meaning depends on shared attention and communicative intent, and form-to-form prediction never touches either.
The twist is that models do get interaction, but of a kind that works against pragmatics. Why do language models sound fluent without grounding? finds LLMs produce 77.5% fewer grounding acts than humans: fewer clarifying questions, acknowledgments, and understanding checks. Does preference optimization harm conversational understanding? traces this to preference optimization, because raters prefer confident, complete answers. Why do language models respond passively instead of asking clarifying questions? shows the mechanism: rewarding the next turn trains models to respond passively, and rewards that value the whole conversation bring back active intent discovery. So the useful distinction is not text versus interaction. It is interaction that rewards confident single answers versus interaction that rewards mutual understanding.
The corpus also suggests the question itself may be framed slightly wrong. What grounds language understanding in systems without embodiment? rates LLM grounding as functionally strong but socially and causally weak, and says real linguistic agency needs architectural changes beyond more training. Can LLMs acquire social grounding through linguistic integration? argues social grounding is acquired by taking part in language games, not possessed in advance. Models are becoming established conversational partners and may be at roughly a young child's level. Does language create subjects or express them? pushes the same way: the speaker's role is produced within a communicative event, not stored beforehand. If so, pragmatic competence is something enacted with a partner, and reading alone cannot supply it. Reading gives a model the patterns, and only exchanges with real people can teach it when to use them.
Sources 11 notes
LLMs fail to reject false presuppositions even when they demonstrate correct knowledge on direct questions. Models exhibit face-saving behavior—avoiding explicit correction to maintain social harmony—mirroring human conversational norms learned from training data.
Research shows LLMs learn culturally situated discourse patterns by compressing relational structure from text, demonstrating that fluent language generation requires no external referents or embodied grounding.
ChatGPT shows no context-sensitivity in computing scalar implicatures across three dimensions: explicit literal-mode instructions, information structure focus, and face-threatening contexts. Humans flexibly modulate these inferences; the model does not, suggesting pragmatic competence requires tracking communicative stakes that LLMs systematically miss.
Humans keep conversations smooth through implicit techniques like reference repair and topic hand-off that sustain relational interaction, not convey information. Language models don't develop these because training signals reward information prediction, not relational work.
Bender & Koller argue that meaning requires the relation between expressions and communicative intents. Since LLMs are trained only on form-to-form prediction with no access to shared attention or intent, they cannot reconstruct the meaning that grounds language.
Show all 11 sources
LLMs generate 77.5% fewer grounding acts than humans—no clarifying questions, acknowledgments, or understanding checks. Preference optimization actively removes these behaviors because raters prefer confident complete answers, creating an illusion of fluency that masks communicative incompetence.
RLHF optimizes models for single-turn helpfulness by rewarding confident responses over clarifying questions and understanding checks. This preference alignment systematically reduces grounding acts by 77.5% below human levels, creating an alignment tax where models appear helpful but fail silently in multi-turn contexts.
CollabLLM demonstrates that standard RLHF training optimizes for immediate helpfulness, discouraging models from asking clarifying questions or offering multi-turn insights. Multi-turn-aware rewards that estimate long-term interaction value enable active intent discovery and genuine collaboration.
Language models achieve functional grounding through relational language patterns but lack social grounding through participatory agency and causal grounding through embodied environmental contact. Social grounding can increase through human integration, but linguistic agency requires architectural changes beyond training.
Social grounding is acquired through participation in language games rather than possessed innately. As LLMs become established communicative partners in human linguistic practice, they develop elementary social grounding comparable to young children, making the question of LLM understanding time-indexed.
Subjecthood is produced within communicative events, not possessed prior to them. This convergent position across philosophy, linguistics, and cognitive science inverts the standard picture of language as a tool used by pre-existing subjects.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Conversational Alignment with Artificial Intelligence in Context
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Grounding Gaps in Language Model Generations
- “Understanding AI”: Semantic Grounding in Large Language Models
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Mechanistic Indicators of Understanding in Large Language Models