SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Do full web pages beat markdown chat for LLM responses?

When LLMs generate complete interactive web pages instead of markdown text, do users prefer them? And how do they compare to pages built by human experts?

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

The paper reports a system where, instead of answering a prompt with markdown text, the LLM generates "a single fully-generated web page and a set of accompanying assets, such as images," rendered as-is in the browser. Against the standard markdown chat baseline, the authors find their generated pages "overwhelmingly preferred by humans," specifically "in 83% of evaluated cases." Measured against PAGEN, a new dataset of pages built by human expert teams for the same prompts, the system's pages are worse overall but "at least comparable in 50% of cases." The authors also report that this capability "is emergent, with substantial improvements from previous models" — older LLMs could not do this reliably, and the jump appeared with the newest generation of models.

The system has three parts: a server exposing tools (image generation, search) that the model can call; a large hand-written system prompt covering goals, planning guidelines, examples, and formatting/tooling instructions; and a set of lightweight post-processors that catch and fix common HTML/CSS/JavaScript errors after generation. The paper frames this as replacing a human product-manager/designer/engineer team with an "instant AI team" assembled per prompt, and argues the result is a shift from a "finite collections of texts" paradigm (fixed apps, fixed templates) to an "infinite catalog" of ephemeral, purpose-built interfaces.

This sits closest to Do generated interfaces outperform text-based chat for most tasks?, which reports a lower preference margin (70%+) for a narrower mechanism: LLM-selected interactive widgets built from a structured, finite-state representation with iterative generation-evaluation refinement. This paper's system instead generates an entire page end to end — HTML/CSS/JS, images, and all — stitched together by prompting and post-processing rather than a structured interface model, and reports a higher preference margin plus two findings the other note lacks: rough parity with human-expert output half the time, and the claim that the whole capability is new to the latest models rather than a steady improvement. It also contrasts with Do generated analysis UIs really work better than chat?, which found generated UIs add authoring rigidity and prompting overhead for analysts composing widgets iteratively; this paper's one-shot, fully generated page sidesteps that composition problem entirely, at the cost of being slow and allowing no user editing after the fact.

The excerpt does not say who the human raters were, how many comparisons the 83% and 50% figures rest on, or what the expert-built PAGEN pages were optimized for, so the comparisons' statistical weight is unclear. It also does not explain what "emergent" means operationally beyond pointing to Tables 3 and 4 the excerpt doesn't include, so the claim that older models simply cannot do this, rather than do it worse, is the authors' characterization rather than something the excerpt demonstrates directly. The paper itself names generation speed (a minute or two per page) and occasional JavaScript/CSS/HTML errors as open limitations, so the implication holds only for settings that can tolerate that latency and residual error rate.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI-assisted work increase total productivity or just shift time? Should GUI agents use structured screen representations instead of end-to-end vision? What prevents LLMs from applying their reasoning knowledge to improve outputs?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 78 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

generative UI full web-page responses are preferred over markdown chat output in 83 percent of cases and emerge only in the newest LLMs