Generative UI: LLMs are Effective UI Generators

Paper · arXiv 2604.09577 · Published February 24, 2026
Knowledge After the Web

AI models excel at creating content, but typically render it with static, predefined interfaces. Specifically, the output of LLMs is often a markdown “wall of text”. Generative UI is a long standing promise, where the model generates not just the content, but the interface itself. Until now, Generative UI was not possible in a robust fashion. We demonstrate that when properly prompted and equipped with the right set of tools, a modern LLM can robustly produce high quality custom UIs for virtually any prompt. When ignoring generation speed, results generated by our implementation are overwhelmingly preferred by humans over the standard LLM markdown output. In fact, while the results generated by our implementation are worse than those crafted by human experts, they are at least comparable in 50% of cases. We show that this ability for robust Generative UI is emergent, with substantial improvements from previous models. We also create and release PAGEN, a novel dataset of expert-crafted results to aid in evaluating Generative UI implementations, as well as the results of our system for future comparisons. Interactive examples can be seen at generativeui.github.io.

Introduction. AI models today generate content: text, code, images, videos, etc. However, the results of these powerful tools are often presented using hard-coded and pre-designed user interfaces. Generative UI is a new modality where the AI model generates not only content, but the entire user experience. This results in custom interactive experiences, including rich formatting, images, maps, audio and even simulations and games, in response to any prompt (instead of the widely adopted “walls-of-text”).

An instant AI team for each prompt. Today, rich visual interfaces exist for common user journeys. Specifically, teams composed of product managers, UX designers, and engineers work for extended periods of time to build amazing rich experiences for broad prompt categories, shared by many users. Generative UI enables us to spin up an instant (AI-based) product management, UX design, and engineering teams, to build an interactive experience, over the course of a minute, for a specific prompt. While not as competent as human experts, Generative UI enables custom experiences for any prompt.

At present, the prevalent UI for interacting with LLMs is a markdown-based chat interface. Specifically, the model outputs markdown (that can include heading, emojis, tables, code-blocks, etc.). These are significantly easier for humans to consume than raw text, yet results created by our Generative UI implementation are overwhelmingly preferred over both (see Table 2). To evaluate our implementation of Generative UI we use human rater preference compared to a set of baselines. We collect and make available PAGEN (see Section 4), a dataset of pages made by human experts for

Related work. The concept of automatically generating user interfaces from high-level descriptions has been a long-standing ambition in Human-Computer Interaction (HCI) and software engineering. Our work builds upon several key areas of research, including natural language interfaces, code generation by large language models (LLMs), and evaluation methodologies for generative systems.

UI Generation from Natural Language Early efforts in this domain often relied on structured inputs or constrained languages to generate interfaces for specific platforms [Puerta et al., 1994, Landay and Myers, 1995]. With the rise of deep learning, approaches evolved to translate visual inputs, such as hand-drawn mockups or screenshots, directly into code [Beltramelli, 2017, Gui et al., 2025]. The recent proliferation of powerful LLMs has enabled the generation of UI code directly from unstructured natural language prompts. Our approach differs by tasking the LLM to generate entire, interactive, and data-driven web applications from a single prompt, effectively acting as an autonomous web developer.

Large Language Models for Code Generation The capabilities of our system are fundamentally enabled by the advancements in code generation by LLMs. This field gained prominence with models like OpenAI’s Codex [Chen et al., 2021], which demonstrated a strong ability to translate natural language into functional code across various languages. Subsequent research has produced a host of powerful code-generating models, such as AlphaCode [Li et al., 2022] and Code Llama [Rozière et al., 2024], that are trained on vast datasets of public code. While these models are often used as assistants for developers (e.g., GitHub Copilot), our work leverages this underlying capability for a different purpose: the autonomous end-to-end generation of a complete user-facing product, not just a code snippet. As we demonstrate, this ability to architect and implement a full application appears to be an emergent property of the most recent state-of-the-art models.

Interaction Paradigms for AI The standard user interface for interacting with LLMs is a chatbased format where the model’s output is rendered as markdown. While an improvement over plain text, this modality is inherently static. Some systems have explored a middle ground, which we term "Templated UI," where an LLM can invoke and populate predefined, interactive widgets from a fixed library to enrich its responses [Pinsky, 2023]. Our work represents a paradigm shift away from both static markdown and constrained templates.

Method. Our Generative UI implementation outputs a single fully-generated web page and a set of accompanying assets, such as images. The page is rendered as-is on the user’s browser. See Figure 2 for a high level overview of the system.

As depicted in Figure 2, we employ 3 main components:

  1. A server exposes several endpoints enabling access to key tools, such as image generation and search. The results can be made accessible to the model (increasing quality) or sent directly to the user’s browser (increasing efficiency). 2. Carefully crafted system instructions. These in turn include: (1) the goal (2) planning and thinking guidelines, (3) examples, and (4) a large set of technical instructions including formatting guidelines, tool endpoints manual, and tips for avoiding common errors. These contribute to the quality of the generated results (see Appendix A.5 for an illustrative prompt from an early research prototype). 3. A set of post-processors. These lightweight components address a set of remaining common issues. Additional post processors deal with error reporting and page analysis. See Appendix A.6.

If desired, our setup allows producing results using a specific style and increased visual consistency across generations. This is done via small changes to the system instructions. Specifically, we experimented with replacing the short “Style” section in our prompt with more detailed variants (which we call “Classic” and “Wizard Green”), specifying colors, fonts, etc. We observe that indeed the generated results follow these styles. Interestingly, the model automatically adapts all elements, including e.g. the generated images and icons to the desired styles. See Figures 3 and 4.

Discussion. We presented a novel implementation of Generative UI, where the model can produce a custom visual interactive interface for any prompt. We show that when ignoring generation speed, our results are overwhelmingly preferred by users over the standard markdown UI (in 83% of evaluated cases, see Table 2). We further show that Generative UI is an emergent capability of the newest and most capable models. As shown in Tables 3 and 4, the use of our newest models results in a significant increase in user preference and a significant reduction in generation errors, vs. previous models. Our implementation relies on a combination of exposing a set of tools in an easy-to-use fashion, detailed system instructions (see Appendix A.5), and a series of post-processors to correct common issues.

PAGEN We created the PAGEN dataset - a curated dataset of expert built web sites for LLM prompts (see Section 4). While the pages created by the expert humans are better than those created by our system, we show that our Generative UI implementation can at least match its quality in 50% of cases. We are making PAGEN available publicly to enable easier evaluation by future research.

Conclusion. A First Step Towards a New Paradigm LLMs transformed the world’s finite collections of texts to an infinite collection, where an ephemeral text is created on the spot for any need. This turned out to be very useful. It is early days for Generative UI, and important limitations exist. Yet, we are excited about a future where users don’t have to pick from a finite library of applications or visual pages, but instead, they have access to an infinite catalog, where the right ephemeral interface is generated on the spot tailored for their need.

Limitations. and Future Directions One primary limitation and an important area for future research is the slow generation speed, which can often take a minute or two. Streaming the generated results allows the users to start interacting with a partially rendered page, reducing this number by about a half. Optimizing the use of techniques such as speculative decoding [Leviathan et al., 2022] could result in further improvements. A second important limitation of Generative UI is that errors (Javascript errors, CSS errors, HTML errors, etc.) can occasionally occur.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can LLMs distinguish between linguistic form and semantic meaning? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? How do network effects and self-selection distort aggregated rating accuracy? Should GUI agents use structured screen representations instead of end-to-end vision? How can agents discover and adapt to user preferences during conversation? Do language models reason through disagreement or only accommodate it? Why do language models fail at sustained therapeutic relationships despite understanding techniques? How does AI adoption reshape collaboration patterns in knowledge work? What design features sustain romantic bonds with AI companion systems? Can AI chatbots provide mental health support without reinforcing harmful beliefs? What are the fundamental limits of prompting for language models? Can AI systems participate in genuine communication or only simulate it? How should humans and AI agents share control and decision-making? What gaps exist between benchmark performance and real deployment outcomes? What prevents LLMs from applying their reasoning knowledge to improve outputs?