Why do LLMs struggle with organizing long-form non-fiction?
Nathan Lambert's textbook writing experience suggests LLMs fail at integrating knowledge across chapters despite handling individual sentences well. The question is whether this reflects a fundamental limitation in how models compress and organize information.
Nathan Lambert, who just finished writing a post-training textbook (Reinforcement Learning from Human Feedback, Manning), argues that LLM progress on long-form non-fiction writing has stalled even as models have gone "from okay to superhuman" at coding and math. He writes that "the models seem genuinely horrible at long-form technical writing" — they can get a sentence or equation right, but asked to write "an entire chapter it'll be a mix of sprinkled with confusing wording, muddled in its organization, and generally a bit off." He used AI extensively for copyediting, LaTeX/TikZ diagrams, and syncing Markdown and LaTeX versions of the book, but says "less than 1%" of the book's explanatory sentences came from a model, included only where he, "as a true expert," judged them as what the reader needed.
Lambert's explanation is that "organizing knowledge is a compression" needed to produce insight, and that today's models do the opposite: they "increase entropy in long-form non-fiction writing." The failure mode he names is compounding error — models can check "every unit of content," a sentence, equation, or figure, but "don't do a good job revisiting components and stringing them together" across many additions, producing what he calls "irreducible compounding errors." He contrasts this with math and code, where he credits RLVR as "a truly magical solution" to the same compounding-error problem — implying long-form writing lacks an equivalent verifiable training signal.
This dovetails with Can LLMs generate more novel ideas than human experts?: Lambert's unit-level competence paired with integration failure is the same generation/evaluation split described there, applied to prose instead of research ideas. It also echoes Do LLM improvements reflect reasoning gains or corpus shifts? in locating the real source of quality in human judgment rather than model reasoning — Lambert says his own "understanding, in the form of intuition, taste, instinct" is what stayed valuable, and that using AI for the writing itself "takes away from that progression." By contrast, where Does polished AI output trick audiences into trusting it? warns that polished AI output can pass as expert judgment to an unwitting audience, Lambert describes the opposite outcome under expert supervision: because he kept "a very close eye" on the few AI sentences he kept, style never substituted for his own judgment.
This is one author's working experience with a small number of named models (GPT-4.5, Kimi K2, Claude) on one book, not a controlled study of writing quality, and Lambert himself is extrapolating a two-to-five-year forecast from it. It doesn't establish why long-form writing specifically lacks a tractable training signal, only his guess that it is "challenging and lacks good training data." The implication he draws, cautiously, is that if models can't yet organize and compellingly present already-established science, claims about them autonomously solving open scientific problems should be treated with skepticism until this narrower, prerequisite skill improves.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What limits language model accuracy in evaluating ideas?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLMs generate more novel ideas than human experts?
Research shows LLM-generated ideas score higher for novelty than expert-generated ones, yet LLMs avoid the evaluative reasoning that characterizes expert thinking. What explains this apparent contradiction?
both find LLMs competent at discrete generation but failing at integrative, structural judgment
-
Do LLM improvements reflect reasoning gains or corpus shifts?
When large language models improve on previously failed tasks, does this show they've learned to reason better, or are they simply reflecting changes in human-written text they train on? Understanding this matters for assessing what LLMs actually know.
both locate the source of writing quality in human judgment, not model reasoning
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
contrasts: under close expert supervision, Lambert kept style from substituting for his own judgment
-
Can LLMs become good editors by learning a writer's taste?
Explores whether large language models can perform well as editors—despite poor creative writing—by being trained on explicit personal taste rubrics rather than generic standards.
extends: teaching an LLM an explicit taste rubric makes it a good editor, matching Lambert's finding that models excel at editing, not long-form writing
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- I wrote an AI textbook — how long until AI can do it better?
- Generalization Bias in Large Language Model Summarization of Scientific Research
- Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
- The Future of AI: Exploring the Potential of Large Concept Models
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Experimental evidence of the effects of large language models versus web search on depth of learning
- AI Meets the Classroom: When Does ChatGPT Harm Learning?
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Original note title
Lambert argues LLM writing ability has stalled on long-form non-fiction because organizing knowledge is a compression the models cannot yet do