Does gaze reveal whether people have achieved common ground?
In collaborative tasks where partners hold different information, does where people look—at the task, at each other, or away—signal moments when they've successfully aligned their understanding? Two studies tested this.
The paper asks "whether gaze provides evidence about grounding" in collaborative tasks where participants hold different private information, and reports that it does, in the same direction across two corpora. In HCRC MapTask, aligned reference interpretations, and in MUNDEX, UND (understood) judgments, are associated with more task-directed gaze and less partner-directed gaze, lower gaze entropy and fewer gaze transitions. The associations are clearest for the participant leading the task: giver-produced references in MapTask and explainer judgments in MUNDEX, the latter also co-varying with the explainee's gaze. Follower-produced references show "near-zero effects."
The mechanism the authors give is the standard grounding account. When information is asymmetric, mutual understanding "cannot be assumed from shared context alone" and must be built and tracked through interaction, citing Clark and Brennan. Gaze is offered as an observable cue to that process, since participants look at task materials, at each other, or away while instructing, checking understanding and coordinating perspectives. The direction of the result echoes earlier map-task observations that partner-directed gaze rises around communicative difficulty. One within-speaker result stands out: in MapTask reference chains, the speaker's gaze entropy is lower at the mention where a previously nonaligned referent becomes aligned, which puts the signal at the moment of repair rather than only across whole conversations.
Against the library, this gives Why do speakers need to actively calibrate shared reference? a nonverbal trace. That note holds that shared surface words guarantee nothing about shared referents; here the evidence that calibration succeeded shows up partly outside the words. It also complicates Why don't conversational AI systems mirror their users' word choices?, which treats alignment as visible in lexical choice; gaze is a second, non-lexical channel of the same alignment work. The predictive gain is modest, which sits well beside Can conversation structure predict dialogue success better than content?: both find signal in how an interaction unfolds rather than what is said, but the gaze features improve only "modestly over controls under grouped cross-validation," and the two studies predict different targets, so the comparison is loose.
The excerpt is silent on sample sizes, effect sizes, causal direction and whether any of this helps a model. It flags its own limits. The two labels "measure different constructs," referential alignment in MapTask against pooled explainee self-reports and explainer judgments in MUNDEX, so the convergence is across related but distinct measures. Which features predict best also differs, with temporal dynamics in MapTask and raw proportions in MUNDEX. The shared partner/task/away categories name gaze targets, not functions: a partner glance may check a landmark reference, and averted gaze has been linked to topic changes. What follows at this strength is narrow: gaze is weak but consistent evidence of grounding in humans, concentrated in whoever leads. A text-only interface has no such channel, so a system there gets grounding evidence only from the words.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does dialogue structure affect linguistic grounding and shared meaning? Should agents decouple planning from perception grounding for better performance?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do speakers need to actively calibrate shared reference?
Explores whether using the same words guarantees speakers mean the same thing. Investigates how referential grounding differs across people and what collaborative work is needed to establish true understanding.
the grounding account this paper supports, adding gaze as an observable trace of calibration in human pairs
-
Why don't conversational AI systems mirror their users' word choices?
Explores whether current dialogue models exhibit lexical entrainment—the human tendency to align vocabulary with conversation partners—and what's needed to bridge this gap in AI communication.
another alignment signal in human dialogue, lexical rather than gaze-based
-
Can conversation structure predict dialogue success better than content?
Does the geometric shape of how dialogue unfolds—timing, repetition, topic drift—matter as much as what people actually say? This explores whether interactive patterns hold signals hidden in word choice alone.
also finds signal in interaction dynamics beyond content, with a different target and stronger reported accuracy
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
- A recipe for annotating grounded clarifications
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- Linguistic Alignment in Conversational AI: A Systematic Review of Cognitive-Linguistic Dimensions, Measurements, and User Outcomes (2020–2025)
- Conversational Alignment with Artificial Intelligence in Context
- Show Me Your Prompts! How Writers Feel About Sharing Prompts in Collaborative Text Editors
- Grounding Gaps in Language Model Generations
Original note title
gaze is evidence of common grounding in both MapTask and MUNDEX, strongest for the participant who leads the task