INQUIRING LINE

After an AI writes website code for you, what actually needs double-checking before you trust it?

What post-processing steps are needed to fix errors in LLM-generated HTML and JavaScript?

This explores what cleanup and checking you need after an LLM writes web code (HTML and JavaScript), and why. The corpus has no paper on HTML or JavaScript repair pipelines, but it says a good deal about which kinds of fixes work on LLM output.


This explores what cleanup and checking you need after an LLM writes web code (HTML and JavaScript), and why. To be clear up front: the collection has no study that lists post-processing steps for generated HTML or JavaScript, such as linting, auto-closing tags or sandboxed execution. What it does have is a set of findings about where LLM errors come from and which corrections work. Taken together, they point to one principle for any cleanup pipeline: errors have to be caught by something outside the model. Asking the model to check its own work isn't enough.

The demand for this is real. Users preferred full LLM-generated web pages over plain markdown chat answers 83% of the time, and those pages matched expert-built ones about half the time, and only with the newest models Do full web pages beat markdown chat for LLM responses?. Read the other way, that means roughly half of generated pages fall short of what an expert would build. Some post-processing is needed no matter how good the model is.

The most useful finding for designing that pipeline is that LLMs can't reliably fix their own reasoning without outside feedback. Asking a model to review and correct its answer, with nothing external to check against, often makes accuracy worse. Having several model instances debate does no better than simple majority voting Can language models fix their own reasoning mistakes?. For web code, this is good news, because code has unusually cheap outside checks. An HTML validator, a JavaScript parser, console errors from a headless browser, or a screenshot comparison can all play the role of the external judge that self-review lacks. One way to frame the whole question: the step that matters most is the one that feeds real error messages back to the model. A 'please double-check your code' prompt adds little. The same reasoning supports calling LLM mistakes 'fabrications' rather than 'hallucinations'. Correct and broken code come from the same generation process, so the fix is a verification layer. Better prompting won't remove the errors Should we call LLM errors hallucinations or fabrications? Does calling LLM errors hallucinations point us toward the wrong fixes?.

The document-editing research suggests what that verification should look for, and the answer depends on the model. Weaker models tend to drop content in obvious ways. Frontier models tend to corrupt it quietly while the surface still looks fine Does model capability change how documents degrade?. Over long edit chains, even top models corrupted about 25% of a document's content, and spot checks missed it Do frontier LLMs silently corrupt documents in long workflows?. Applied to web code, a page that validates cleanly can still have a broken event handler, a deleted section, or a wrong value buried in a script. So a syntax check isn't enough. You also need behavior checks (does the button still work?) and diffs against the previous version when a model edits existing code. Giving the model better editing tools doesn't solve this, because the errors come from its judgment about what to change, not from clumsy editing tools Can better tools fix LLM document editing errors?.

Finally, some errors happen before any code is written. In conversations where requirements arrive bit by bit, models commit early to wrong assumptions and rarely recover, losing about 39% of their performance Why do language models fail in gradually revealed conversations?. For iterative web building, the practical lesson is that restarting with a single, fully stated spec often beats a long run of 'now also make it do X' patches. Work on turning LLMs into agents that take actions makes the broader point: reliability comes from the system built around the model, not from the model alone Can you turn an LLM into an agent by just fine-tuning?. The conclusion for web code is that post-processing is not a cleanup step bolted on at the end. It is the system that decides whether the generated page actually works.


Sources 9 notes

Do full web pages beat markdown chat for LLM responses?

Users strongly prefer LLM-generated full web pages over markdown replies, with 83% preference in direct comparisons. Generated pages match expert-built pages in quality roughly half the time, and this capability appears only in the newest models.

Can language models fix their own reasoning mistakes?

Across GPT-3.5, GPT-4, GPT-4-Turbo, and Llama-2, self-correction without external labels degrades reasoning accuracy. Multi-agent debate gains match plain self-consistency at identical cost, suggesting debate is consistency voting, not genuine correction.

Should we call LLM errors hallucinations or fabrications?

LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.

Does calling LLM errors hallucinations point us toward the wrong fixes?

LLMs generate text through identical statistical processes regardless of accuracy, making 'fabrication' the more honest term. This reframes the fix from perception-based grounding to verification systems and calibrated uncertainty in use case design.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Show all 9 sources
Do frontier LLMs silently corrupt documents in long workflows?

Even the strongest models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) degrade documents by ~25% over long relay workflows across 52 domains. Degradation decelerates but never plateaus, and errors compound silently, remaining undetected in spot-checked outputs.

Can better tools fix LLM document editing errors?

DELEGATE-52 shows that agentic tool access fails to improve performance on long-horizon document tasks. The degradation mechanism originates upstream in the model's judgment about what to change, not in editing interface limitations.

Why do language models fail in gradually revealed conversations?

Across 200,000+ conversations, all major LLMs show 39% average performance drop in multi-turn settings due to locking into incorrect early guesses. Agent mitigations recover only 15-20% of this loss.

Can you turn an LLM into an agent by just fine-tuning?

Converting LLMs to action-capable systems requires four distinct stages: curating action-environment-user datasets, training for action grounding, integrating agent infrastructure with memory and tools, and rigorous safety evaluation. The surrounding system and harness determine whether actions are grounded or hallucinated.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.