Can LLMs become good editors by learning a writer's taste?
Explores whether large language models can perform well as editors—despite poor creative writing—by being trained on explicit personal taste rubrics rather than generic standards.
Jasmine Sun argues that large language models remain poor creative writers but can become good editors — a split she traces to three causes, not one. Even AI optimists concede the writing gap: Sam Altman guesses GPT-7 might manage only "a real poet's okay poem," and Tyler Cowen and Patrick Collison's "New Aesthetics" grant notes "we haven't seen much great work that only uses AI." Sun's own editing tool, built by teaching Claude her personal rubric, closed part of that gap on the editing side: "The resulting tool is as good as many human editors I've had."
Sun's mechanism has three parts. "Good writing" is hard to train because it's hard to evaluate — current writing evals are "mostly dumb," she reports, citing a labeler asked to grade fanfiction on "factuality." It is also not a business priority, so other post-training goals crowd it out, leaving chatbots with "factory settings designed for a sycophantic corporate assistant, not a creative genius." And it requires grounding in lived experience that she holds nonfiction in particular needs. Editing, she argues, asks less of this: a good editor needs curiosity and taste more than a distinct voice, which is "closer to a role that LLMs can play" — provided the model is taught whose taste to apply. She cites Tuhin Chakrabarty's 2023 rubric study, where human and LLM judges reached high agreement scoring short stories against 14 literary criteria, as evidence that "literary quality might be measured — and thus trained for." Her own practice generalizes this: teaching Claude her specific voice and goals made its feedback track her writing rather than impose generic rules, so it stopped flagging her slang as "too casual" and started flagging thin reporting instead of inventing scenes.
This complicates the flat claim in Can LLMs generate more novel ideas than human experts? that LLMs cannot perform evaluative stance-taking. Sun's case is narrower and more hopeful: evaluation is learnable when an explicit, expert-built rubric substitutes for the taste a model lacks by default — the gap is default taste, not evaluative capacity as such. That reframes Does polished AI output trick audiences into trusting it? and Can language models truly understand literary style?: a model that only detects style patterns without taste will, by Sun's account, default to "rigid standards" — exactly the generic editing-app behavior she contrasts with her personalized tool. It also sits against Does polished writing actually signal better quality work?: Sun's worry runs the other direction, that an untaught model's taste is bad polish (sycophantic, rule-bound) rather than deceptively good polish.
The excerpt does not establish that personalized-rubric editing generalizes beyond Sun's own case — "it worked" describes one practitioner's workflow, not a controlled comparison against the "variously disappointing editing apps" she names only in passing. Nor does the Chakrabarty finding of judge agreement on short-story rubrics demonstrate that the same approach scales to nonfiction editing, Sun's own domain; she imports the result by analogy. The evidence supports a narrower claim than "LLMs can learn taste": that an explicit, person-specific rubric measurably improves an LLM's editorial fit for that one person, while generative creative quality remains unaddressed by the same fix.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can readers reliably distinguish AI-written text from human writing? What limits language model accuracy in evaluating ideas? How can we detect and account for LLM involvement in academic writing?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLMs generate more novel ideas than human experts?
Research shows LLM-generated ideas score higher for novelty than expert-generated ones, yet LLMs avoid the evaluative reasoning that characterizes expert thinking. What explains this apparent contradiction?
Sun's taught-rubric editing complicates the claim that evaluative stance-taking is structurally absent from LLMs
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
Sun's "factory settings" taste gap is why untaught models default to presentation-as-proxy rather than genuine judgment
-
Can language models truly understand literary style?
LLMs detect stylistic patterns with high accuracy, but can they grasp why those patterns matter? This explores the gap between surface-level pattern recognition and meaningful interpretation.
pattern-only style detection is what Sun says generic editing apps rely on before being taught an actual rubric
-
Does polished writing actually signal better quality work?
When evaluators judge applications and manuscripts, does rhetorical sophistication predict merit, or does it distract from verifiable evidence of competence and rigor?
contrasts the direction of the taste failure: Sun frames untaught AI taste as bad and rule-bound, not deceptively convincing
-
Why do LLMs struggle with organizing long-form non-fiction?
Nathan Lambert's textbook writing experience suggests LLMs fail at integrating knowledge across chapters despite handling individual sentences well. The question is whether this reflects a fundamental limitation in how models compress and organize information.
Extends A: Lambert's finding that LLMs aid editing but can't yet organize long-form structure specifies why LLMs stay bad writers
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- why LLMs are bad writers but good editors
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Measuring AI "Slop" in Text
- GhostWriter: Augmenting Collaborative Human-AI Writing Experiences Through Personalization and Agency
- Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing?
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
- Scientific production in the era of Large Language Models
Original note title
Sun argues teaching an LLM a writer's own taste rubric turns it into a good editor even though it remains a bad creative writer