Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee’s gaze. In same-speaker MapTask reference chains, the speaker’s gaze entropy is lower at the mention where a previously nonaligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX.
Introduction. In collaborative tasks where participants hold different private information, mutual understanding cannot be assumed from shared context alone. It must be built and tracked through interaction (Clark and Wilkes-Gibbs, 1986; Clark and Brennan, 1991). Gaze is an observable cue to this process: participants look at task materials, at each other, or away while giving instructions, checking understanding, and coordinating their perspectives. Some corpora annotate gaze from video as discrete categories of where participants look, rather than as eye-tracking coordinates. These annotations can be used to study the relationship between gaze and grounding, but they are often corpus-specific, making it difficult to compare across tasks. We study two settings where information asymmetry forces participants to continuously coordinate understanding. In HCRC MapTask (Anderson et al., 1991), a giver and a follower navigate with maps that differ in their landmarks; perspectivist grounding labels record each participant’s interpretation separately (Li et al., 2026a).
Discussion / Conclusion. Shared categories, task-specific meanings The direction of these associations is the same in both corpora, echoing map-task observations that partner-directed gaze increases around communicative difficulty (Boyle et al., 1994; Nakano et al., 2003; Murat and Vogel, 2026). The two labels measure different constructs: MapTask records referential alignment, whereas MUNDEX pools explainees’ self-reports and explainers’ judgments, so the convergence spans related but distinct grounding measures. Which features carry predictive signal differs: temporal dynamics score highest in MapTask and raw proportions in MUNDEX. The shared categories also name gaze targets rather than functions. In MapTask, a partner glance may check a landmark reference; in MUNDEX, gaze averted from the partner has also been linked to topic changes (Lazarov and Grimminger, 2026), so it Gaze and interactional role Significant associations concentrate in giver-produced references and explainer judgments, whereas follower-produced references show near-zero effects.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do transformer attention mechanisms implement memory and algorithmic functions? Why should disagreement be treated as signal in collaborative reasoning? Does conversational format create illusions of genuine AI communication? How can LLM user simulators model realistic goal-driven conversation? How do language models establish social grounding in human dialogue?- Does functional grounding through discourse patterns count as genuine semantic meaning?
- How do humans maintain separate mental contexts during a single conversation?
- Does social grounding in language improve through iterative human integration?
- What makes social grounding different from constitutive linguistic agency?
- Can convention formation improve communicative grounding beyond word sharing?
- How does speaker responsibility shape whether something counts as communication?
- How does linguistic coordination build shared reference between conversational partners?
- How does shared reference and grounding affect assumption detection in dialogue?
- What role do first-person pronouns play in sustaining collaborative conversation tone?
- Why can't static grounding alone close the gap between agreement and understanding?
- What role does dynamic grounding play in achieving real mutual understanding?
- What is the difference between static and dynamic grounding in dialogue?