Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

Paper · arXiv 2609.18011 · Published September 16, 2026
NLP and Linguistics

In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee’s gaze. In same-speaker MapTask reference chains, the speaker’s gaze entropy is lower at the mention where a previously nonaligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX.

Introduction. In collaborative tasks where participants hold different private information, mutual understanding cannot be assumed from shared context alone. It must be built and tracked through interaction (Clark and Wilkes-Gibbs, 1986; Clark and Brennan, 1991). Gaze is an observable cue to this process: participants look at task materials, at each other, or away while giving instructions, checking understanding, and coordinating their perspectives. Some corpora annotate gaze from video as discrete categories of where participants look, rather than as eye-tracking coordinates. These annotations can be used to study the relationship between gaze and grounding, but they are often corpus-specific, making it difficult to compare across tasks. We study two settings where information asymmetry forces participants to continuously coordinate understanding. In HCRC MapTask (Anderson et al., 1991), a giver and a follower navigate with maps that differ in their landmarks; perspectivist grounding labels record each participant’s interpretation separately (Li et al., 2026a).

Discussion / Conclusion. Shared categories, task-specific meanings The direction of these associations is the same in both corpora, echoing map-task observations that partner-directed gaze increases around communicative difficulty (Boyle et al., 1994; Nakano et al., 2003; Murat and Vogel, 2026). The two labels measure different constructs: MapTask records referential alignment, whereas MUNDEX pools explainees’ self-reports and explainers’ judgments, so the convergence spans related but distinct grounding measures. Which features carry predictive signal differs: temporal dynamics score highest in MapTask and raw proportions in MUNDEX. The shared categories also name gaze targets rather than functions. In MapTask, a partner glance may check a landmark reference; in MUNDEX, gaze averted from the partner has also been linked to topic changes (Lazarov and Grimminger, 2026), so it Gaze and interactional role Significant associations concentrate in giver-produced references and explainer judgments, whereas follower-produced references show near-zero effects.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do transformer attention mechanisms implement memory and algorithmic functions? Why should disagreement be treated as signal in collaborative reasoning? Does conversational format create illusions of genuine AI communication? How can LLM user simulators model realistic goal-driven conversation? How do language models establish social grounding in human dialogue? How can AI alignment serve diverse human preferences at scale? What makes dialogue-based explanation more successful than monologue? What distinguishes dynamic from static grounding in dialogue systems? Is embodied interaction necessary for language meaning and genuine agency? How do interface design choices shape consciousness attribution? How can emotions function as reliable information in reasoning and cognitive systems? Can AI-generated outputs constitute genuine knowledge or valid claims? How do neural networks separate factual knowledge from reasoning abilities? Does RLHF training sacrifice accuracy and grounding for user agreement?