Line of inquiry
Inquiring lines›How do language models construct a…›How do dialogue systems achieve ge…›this line of inquiry
Should GUI agents use structured representations instead of raw pixels?
A broader line of inquiry — a family of 24 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 24
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does explicit screen parsing outperform pure vision in GUI agents?
- Should GUI agents use intermediate structured representations instead of raw pixels?
- Can specialized perception components replace end-to-end vision in GUI agents?
- Can screen perception be effectively decoupled from planning in GUI agents?
- Why does identifying UI element types and locations enable downstream task learning?
- What role does visual perception play alongside accessibility tree information?
- Can text-based and vision-based screen understanding achieve similar performance?
- Can parsing screens into structured elements before acting improve vision models?
- Should GUI perception happen inside or outside the foundation model?
- How does UI-guided token selection reduce compute compared to standard vision?
- Why does pure-vision underperform when parsing semantics and action prediction mix?
- Why do small specialized models match frontier multimodal models on screen tasks?
- What makes high-quality GUI instruction data different from general vision data?
- Why do GUI agents need pixels while document systems can use bounding boxes?
- What makes accessibility trees insufficient compared to visual GUI understanding?
- Why do static screenshot models fail to capture multi-step UI task intent?
- How does annotation-based pretraining compare to self-supervised video masking for screen understanding?
- How do agents parse HTML differently than human browsers render it?
- Why do multimodal chatbots fail at GUI element grounding tasks?
- What design discipline replaces navigation and layout in AI systems?
- What visual patterns transfer between infographic and UI tasks when trained jointly?
- How does serializing screen layout to text preserve spatial relationships?
- What temporal signals in screen recordings matter most for task understanding?
- What document layouts benefit most from bounding box representations?