An AI interface rarely removes thinking; it moves it, deciding whether you imagine, choose, or judge the output.
How do interface designs shape what cognitive work users actually perform?
This explores how the shape of an AI interface (open chat box, generated dashboard, guided options) changes which mental tasks fall to the user, such as imagining, specifying, evaluating or tracking context, rather than just how pleasant the interface feels.
This explores how an interface decides which thinking the user has to do, and which thinking it quietly takes off their plate. The corpus suggests interfaces rarely remove cognitive work. They move it somewhere else. A blank chat box asks the user to envision: to know what they want and say it clearly before they've seen anything. Research on the 'gulf of envisioning' finds this is often impossible, because intent develops through interaction rather than existing beforehand. Interfaces that show model-generated options change the task from open-ended imagining to choosing among options, which is a much easier kind of work Why can't users articulate what they want from AI?.
Generated interfaces push this further. When an LLM builds a task-specific UI, such as a dashboard, a set of controls or an animation, instead of returning a block of text, users prefer it in over 70% of cases. Structure seems to do the comprehension work that prose leaves to the reader Do generated interfaces outperform text-based chat for most tasks?. But the work doesn't disappear. A study of generated analysis tools found that widgets made results clearer but harder to change halfway through a task. Customizing them required something close to engineering-style specification from people who aren't programmers Do generated analysis UIs really work better than chat?. The trade-off looks built in: the easier a UI is to read, the more thinking it has already locked in for you.
Conversational design shifts work in a less visible way. Chat interfaces draw on skills people have built over a lifetime of talking with other people: turn-taking, shared assumptions, repair when something goes wrong. The system on the other end doesn't actually communicate in that sense. So when the interaction breaks down, it feels like user error, but the cause is in the design Why do users fail with AI interfaces designed like conversations?. There is also a hidden tracking burden. In conventional software, the user can learn a stable layout. AI context (the prompt, the conversation history, retrieved data, hidden state) keeps changing, so the user can never fully keep track of what the system is currently working from How does AI context differ from conventional software context?. Meanwhile, users judge these systems mostly on perceived competence, which accounts for about half of how they model a dialogue partner How do users mentally model dialogue agent partners?. A smooth conversational surface can therefore win trust that its actual behavior hasn't earned.
The less obvious finding comes from AI agents that operate software. When GPT-4V had to read raw screenshots, it failed partly because it was doing two jobs at once: working out what each icon means and deciding what to click. Pre-parsing the screen into labeled elements let it focus on the action alone Why do vision-only GUI agents struggle with screen interpretation?. Agent S got similar gains by separating planning from pinpointing the exact element to act on, using accessibility trees (the structured map of on-screen elements that operating systems keep for screen readers) Can structured interfaces help language models control GUIs better?. Humans and models show the same pattern: an interface that bundles interpreting what's on screen with deciding what to do overloads whoever uses it, and splitting the two helps.
One more point: interfaces don't only hand out cognitive work, they can also watch it happen. Signals like gaze, hesitation and typing speed can let a system infer a user's mental state and time its help so it doesn't interrupt their flow. The same signals could also be used for manipulative profiling Can AI systems read cognitive state from interaction patterns alone?. So when you ask what work an interface makes you do, it's also worth asking what that work reveals about you.
Sources 9 notes
Intent develops through interaction, not in isolation. Since AI models respond rather than probe, they miss opportunities to help users discover unarticulated requirements. Structured dialogue that presents model-generated options shifts the cognitive burden from open-ended envisioning to constrained evaluation.
Research shows users strongly prefer LLM-generated interactive interfaces—dashboards, tools, animations—over text blocks, especially for structured and information-dense tasks. Structured representation and iterative refinement reduce cognitive load.
TaskArtisan found that GUI widgets improve clarity and presentation in LLM-assisted analysis but introduce rigidity and prompting overhead. This trade-off between malleability and specification appears unavoidable: easier-to-use UIs are harder to customize mid-workflow, while flexible UIs demand engineering-style thinking from non-programmers.
AI interfaces that use conversational design conventions trigger users' lifelong communication skills, but AI doesn't actually communicate. This mismatch causes interaction failures that feel like user error but originate in design.
AI interactions operate on a substrate of constantly shifting context—prompt, history, retrieved data, hidden state—that users cannot internalize like traditional UIs. This structural mutability demands a new design discipline centered on context engineering rather than interface design.
Show all 9 sources
The Partner Modelling Questionnaire reveals that perceived competence dominates user impressions (49% of variance), followed by human-likeness (32%) and communicative flexibility (19%). This three-factor structure reflects how people evaluate dialogue partners against both functional and social standards.
OmniParser demonstrates that GPT-4V fails when forced to simultaneously identify icon meanings and predict actions from raw screenshots. Pre-parsing screenshots into structured semantic elements with descriptions lets the model focus solely on action prediction, removing the composite-task bottleneck.
Agent S's dual-input design—visual input for environmental understanding plus image-augmented accessibility trees for grounding—achieved 9.37% improvement over baseline by factoring planning and grounding into separate optimization paths rather than forcing end-to-end prediction.
Research shows AI systems can instrument multimodal behavioral signals (gaze, hesitation, speed) to read cognitive state during interaction, preserving flow by avoiding disruptive explicit probes. However, the same substrate enables both helpful timing and manipulative profiling.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Bridging the gulf of envisioning: Cognitive design challenges in llm interfaces.
- Generative Interfaces for Language Models
- MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
- Design Theater: Evaluating the Gap Between User-Facing Design Reasoning and Implementation in Generative UI Tools
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent
- TaskArtisan: Designing Composable Generative Widgets for LLM-Assisted Analysis
- OmniParser for Pure Vision Based GUI Agent
- Generative UI: LLMs are Effective UI Generators