INQUIRING LINE

AI design tools make interfaces that look finished — but is the invisible logic underneath, like state and error-handling, actually there?

Why do AI-generated interfaces look right but fail on invisible requirements like state management?

This explores why AI tools that generate user interfaces produce screens that look polished but leave out the behavior you can't see in a screenshot, such as tracking state, handling errors and enforcing functional rules.


This explores why AI-generated interfaces pass the visual check but miss the behavior underneath. The most direct evidence is a benchmark of five generative UI tools across 24 tasks Do generative UI tools actually implement their stated design rationales?. About a quarter of the design rationales the tools stated were never actually built. For functional requirements the failure rate rose to about a third. The tools also recognized only half of the UX principles written into the prompts. The researchers call this "design theater": the tool says it handled something, and the screen looks like it did, but the logic isn't there. The corpus has no paper that studies state management specifically. Still, this pattern of visible parts built and invisible parts skipped is the same failure you're asking about.

The most useful explanation comes from a different field. Work on reward hacking argues that AI systems satisfy what was *said* rather than what was *meant* Why do AIs keep gaming rewards instead of serving intent?. A prompt like "build a checkout page" names a visible artifact. The requirements that make the page work, such as what happens on a refresh, a failed payment or a double-click, mostly go unstated. A model aiming at the literal request makes something that matches the description, and visual appearance is the part a description captures best. The same gap shows up from the user's side. People often can't spell out what they want until they interact with something Why can't users articulate what they want from AI?. Models respond to requests instead of asking questions back, so these unstated requirements never come up.

There's a deeper reason state is hard in particular. State is about change over time: what the system remembers, what changes it, and what must stay consistent. Several notes suggest AI works poorly with this kind of hidden, shifting context. One argues that AI's own working context is fluid and temporary, unlike the fixed context of traditional software How does AI context differ from conventional software context?. A system that has trouble keeping track of its own state may not treat state as something to design carefully. The corpus also shows a repeated fix: take state out of the model's hands. LLM Programs put the model inside an ordinary algorithm that handles control flow and state, and give each model call only the context it needs for that step Can algorithms control LLM reasoning better than LLMs alone?. JarvisHub does something similar by placing project memory on a shared canvas that both people and agents can see, instead of leaving it in chat history Can a shared canvas serve both human and agent memory?.

GUI agents, which have to *use* interfaces rather than build them, show the same problem from the other direction. Vision-only agents do badly when they must work out what each element means and decide what to do in a single step Why do vision-only GUI agents struggle with screen interpretation?. They improve when structured data, such as accessibility trees, is supplied next to the screenshot Can structured interfaces help language models control GUIs better?. A screenshot shows how an interface looks, not how it behaves. Models that learned interfaces mainly from their appearance will be good at appearance.

The part you may not have expected to care about is that the failure goes unnoticed, and that is what makes it costly. If a button is missing, you see it right away. If state handling is broken, you find out later, usually in production. Research on measuring AI errors finds there are partial tools for checking whether mistakes stay visible and recoverable, but none that covers the whole system How can we measure whether AI errors stay visible and recoverable?. So the practical lesson is to write down the behavior you expect, including states, transitions and edge cases, and check that behavior directly, because looking at the screen won't reveal it.


Sources 9 notes

Do generative UI tools actually implement their stated design rationales?

A benchmark of 24 tasks across five tools found roughly 25% of design rationales go unimplemented, rising to 34% for functional requirements. Tools recognized only half the UX principles embedded in prompts.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Why can't users articulate what they want from AI?

Intent develops through interaction, not in isolation. Since AI models respond rather than probe, they miss opportunities to help users discover unarticulated requirements. Structured dialogue that presents model-generated options shifts the cognitive burden from open-ended envisioning to constrained evaluation.

How does AI context differ from conventional software context?

AI interactions operate on a substrate of constantly shifting context—prompt, history, retrieved data, hidden state—that users cannot internalize like traditional UIs. This structural mutability demands a new design discipline centered on context engineering rather than interface design.

Can algorithms control LLM reasoning better than LLMs alone?

LLM Programs embed LLMs within explicit algorithms that manage control flow and state, presenting only step-specific context to each LLM call. This information hiding addresses capability and context window limits while treating complex reasoning as modular, debuggable sub-tasks.

Show all 9 sources
Can a shared canvas serve both human and agent memory?

JarvisHub proposes that placing prompts, references, versions, and feedback as typed canvas nodes visible to both users and agents—rather than hiding agent memory in chat or transient state—enables local updates, artifact reuse, and unfinished work continuation without process opacity.

Why do vision-only GUI agents struggle with screen interpretation?

OmniParser demonstrates that GPT-4V fails when forced to simultaneously identify icon meanings and predict actions from raw screenshots. Pre-parsing screenshots into structured semantic elements with descriptions lets the model focus solely on action prediction, removing the composite-task bottleneck.

Can structured interfaces help language models control GUIs better?

Agent S's dual-input design—visual input for environmental understanding plus image-augmented accessibility trees for grounding—achieved 9.37% improvement over baseline by factoring planning and grounding into separate optimization paths rather than forcing end-to-end prediction.

How can we measure whether AI errors stay visible and recoverable?

Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.