Do spatially-anchored previews stabilize VR geometry editing tasks?
This study explores whether showing users visual previews of candidate interpretations anchored in the 3D workspace helps them reach task goals more smoothly and with fewer conversation turns when working with an LLM-assisted geometry editor.
The paper asks whether dialogue-based disambiguation can be "augmented with spatially-anchored graphical previews" when a user gives an ambiguous command to an LLM that edits a VR scene through numerical parameters. In a within-subjects study with 24 participants doing complex geometry editing tasks, the abstract reports that a hybrid of clarification questions and graphical previews "can support better interaction stability with fewer conversation rounds while improving user experience," relative to a condition with no disambiguation. The discussion narrows this: clarification questions alone (CQ) and the combined condition (CQGP) both significantly reduced task progression variability, measured by MSSD, and CQGP "required significantly fewer conversation rounds to complete the task."
The stability result is the one that matters, because the performance result is flat. The maximum closeness score "did not differ significantly across conditions," so disambiguation did not raise the best state users reached. What changed was the path: less erratic progression toward it, and in the hybrid condition fewer conversational turns spent getting there. The paper frames this as text and visual feedback guiding users "towards a smoother task progression trajectory," which it calls novel against 2D interfaces, where disambiguation improved task performance. Its motivation is the "limited user agency due to ambiguous user input" in LLM-assisted scene editing. The previews are a hybrid of 2D UI image previews and spatially-anchored 3D overlays, so the user sees candidate interpretations in the scene rather than reading them in a question.
Against the vault, this is a spatial variant of a text-side result. Which clarifying questions actually improve user satisfaction? concerns what a good clarifying question looks like in text, and here CQ alone already improves stability, which is consistent with clarification helping. The added claim is that a preview anchored in the workspace lets the hybrid condition resolve the ambiguity in fewer rounds. Why do language models fail in gradually revealed conversations? describes the failure that ambiguity produces in text-only multi-turn use, and this study looks at a related setting with a different remedy: show the user the candidate outcome. The paper also states its scope: it "does not study disambiguation on the input side," unlike the speech-and-pointing work on referential disambiguation. That separates it from Can reference resolution work as a language modeling problem?, which works on the input side.
The excerpt leaves several things open. It names three conditions (no disambiguation, CQ, CQGP), so the contribution of previews alone cannot be isolated, and it does not say whether CQGP differs from CQ on stability. It does not state the baseline for the "fewer conversation rounds" comparison, effect sizes, task details or the qualitative themes behind "improving user experience." It gives no evidence on why previews help, and it does not test whether users' intent matured, the concern in How do users actually form intent when prompting AI systems?. A cautious reading is that, in this one task family and sample, showing a candidate result in the scene reduced conversational effort and erratic progression without lifting the best outcome.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Should agents decouple planning from perception grounding for better performance? Why do some clarifying approaches produce understanding while others just satisfy? What prevents conversational agents from taking initiative in dialogue?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Which clarifying questions actually improve user satisfaction?
Not all clarification helps equally. This explores whether asking users to rephrase their needs works as well as asking targeted questions about specific information gaps.
text-side evidence that clarification design matters; here clarification alone also steadies progression in a VR editing task.
-
Why do language models fail in gradually revealed conversations?
Explores why LLMs perform 39% worse when instructions arrive incrementally rather than upfront, and whether they can recover from early mistakes in multi-turn dialogue.
the text-only multi-turn failure that ambiguity produces; this study tests a visual remedy in a spatial setting.
-
Can reference resolution work as a language modeling problem?
Can conversational, background, and on-screen references be resolved by converting them into text and using language models instead of specialized multimodal systems? This matters because it could enable efficient, on-device reference understanding.
input-side referential resolution, which this paper explicitly leaves out of scope.
-
How do users actually form intent when prompting AI systems?
Users face a 'gulf of envisioning'—they must simultaneously imagine possibilities and express them to language models. This cognitive gap creates breakdowns not from AI incapability but from users struggling to articulate what they truly need.
frames the underlying problem; this study measures progression stability, not intent maturation.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Beyond Conversations: Spatially-Anchored Previews for Intent Disambiguation in LLM-Assisted Geometry Editing in Virtual Reality
- Harnessing LLMs Without Surrendering Control: Delegation Boundaries in Visual Data Storytelling Authoring
- MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
- LLMs Get Lost In Multi-Turn Conversation
- MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
- A recipe for annotating grounded clarifications
- Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations
- Bridging the gulf of envisioning: Cognitive design challenges in llm interfaces.
Original note title
clarification questions plus spatially-anchored previews give more stable task progression with fewer conversation rounds in LLM-assisted VR geometry editing