Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do architectural choices affec…›this line of inquiry
Should GUI agents use structured screen representations instead of end-to-end vision?
A broader line of inquiry — a family of 53 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 53
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What tensions emerge when AI models generate interfaces instead of rule-based systems?
- Should GUI agents use intermediate structured representations instead of raw pixels?
- Do users notice when generative interfaces don't match their own stated design principles?
- Can specialized perception components replace end-to-end vision in GUI agents?
- Should GUI perception happen inside or outside the foundation model?
- Can screen perception be effectively decoupled from planning in GUI agents?
- Why does explicit screen parsing outperform pure vision in GUI agents?
- Do dynamically generated interfaces perform better than pre-built ones?
- Why does identifying UI element types and locations enable downstream task learning?
- What makes high-quality GUI instruction data different from general vision data?
- Can API-first interaction replace traditional UI-based agent interfaces?
- Can generative UIs maintain consistency over time without becoming rigid to users?
- Can generative interfaces help users articulate what they actually want?
- Can traditional UX methods work for autonomous AI systems?
- Should AI interfaces keep manual GUI controls as a fallback?
- Can interface design alone overcome lack of awareness about AI tool capabilities?
- How does UI-guided token selection reduce compute compared to standard vision?
- Do specialized interfaces outperform generic chatbots for domain-specific work?
- What design discipline replaces navigation and layout in AI systems?
- How do interface designs shape what cognitive work users actually perform?
- How do typed adapters move learned changes between isolated surfaces?
- Can designers hide AI context complexity behind a stable user interface?
- Why do AI-generated interfaces look right but fail on invisible requirements like state management?
- Why do traditional interfaces bypass the intention formation problem that language models expose?
- Why does pure-vision underperform when parsing semantics and action prediction mix?
- What role does visual perception play alongside accessibility tree information?
- Should agents use APIs or GUI interaction for efficiency?
- How does API-first interaction compare to generative interface approaches?
- Can LLM-generated pages achieve quality parity with expert-designed interfaces at scale?
- Why do dynamic UIs reduce cognitive load but complicate user control and predictability?
- Why do multimodal chatbots fail at GUI element grounding tasks?
- Can natural language help users modify widget composition during analysis work?
- How do generated interfaces compare to chat when tasks require workflow changes?
- How do agents perceive and traverse typed node-and-link structures on a canvas?
- Why do static screenshot models fail to capture multi-step UI task intent?
- What makes accessibility trees insufficient compared to visual GUI understanding?
- How do generated UI capabilities differ between older and newer LLM models?
- Do task-specific interfaces outperform conversational chat in practical settings?
- Why do GUI agents need pixels while document systems can use bounding boxes?
- How do agents parse HTML differently than human browsers render it?
- What trade-offs exist between one-shot full page generation and iterative widget composition?
- What emergent abilities appear only in truly unified multimodal systems?
- How can analysts customize generated UIs without learning to think like engineers?
- Why do APIs outperform UIs for agent task completion?
- Can multimodal architectures successfully integrate vision without replicating past failures?
- What evidence shows canvas workspaces recover from failures better than chat baselines?
- What types of tasks benefit most from dynamically generated interfaces?
- What makes some analysis tasks stable enough for rigid generated interfaces?
- Why did product managers gain more from Figma Make than professional designers?
- What visual patterns transfer between infographic and UI tasks when trained jointly?
- How does serializing screen layout to text preserve spatial relationships?
- How do malleable software and adaptive UI differ in their approach to change?
- What document layouts benefit most from bounding box representations?