SYNTHESIS NOTE
Topics›AI at Work›this note

Can LLMs become good editors by learning a writer's taste?

Explores whether large language models can perform well as editors—despite poor creative writing—by being trained on explicit personal taste rubrics rather than generic standards.

Synthesis note · 2026-10-09 · sourced from AI at Work

Jasmine Sun argues that large language models remain poor creative writers but can become good editors — a split she traces to three causes, not one. Even AI optimists concede the writing gap: Sam Altman guesses GPT-7 might manage only "a real poet's okay poem," and Tyler Cowen and Patrick Collison's "New Aesthetics" grant notes "we haven't seen much great work that only uses AI." Sun's own editing tool, built by teaching Claude her personal rubric, closed part of that gap on the editing side: "The resulting tool is as good as many human editors I've had."

Sun's mechanism has three parts. "Good writing" is hard to train because it's hard to evaluate — current writing evals are "mostly dumb," she reports, citing a labeler asked to grade fanfiction on "factuality." It is also not a business priority, so other post-training goals crowd it out, leaving chatbots with "factory settings designed for a sycophantic corporate assistant, not a creative genius." And it requires grounding in lived experience that she holds nonfiction in particular needs. Editing, she argues, asks less of this: a good editor needs curiosity and taste more than a distinct voice, which is "closer to a role that LLMs can play" — provided the model is taught whose taste to apply. She cites Tuhin Chakrabarty's 2023 rubric study, where human and LLM judges reached high agreement scoring short stories against 14 literary criteria, as evidence that "literary quality might be measured — and thus trained for." Her own practice generalizes this: teaching Claude her specific voice and goals made its feedback track her writing rather than impose generic rules, so it stopped flagging her slang as "too casual" and started flagging thin reporting instead of inventing scenes.

This complicates the flat claim in Can LLMs generate more novel ideas than human experts? that LLMs cannot perform evaluative stance-taking. Sun's case is narrower and more hopeful: evaluation is learnable when an explicit, expert-built rubric substitutes for the taste a model lacks by default — the gap is default taste, not evaluative capacity as such. That reframes Does polished AI output trick audiences into trusting it? and Can language models truly understand literary style?: a model that only detects style patterns without taste will, by Sun's account, default to "rigid standards" — exactly the generic editing-app behavior she contrasts with her personalized tool. It also sits against Does polished writing actually signal better quality work?: Sun's worry runs the other direction, that an untaught model's taste is bad polish (sycophantic, rule-bound) rather than deceptively good polish.

The excerpt does not establish that personalized-rubric editing generalizes beyond Sun's own case — "it worked" describes one practitioner's workflow, not a controlled comparison against the "variously disappointing editing apps" she names only in passing. Nor does the Chakrabarty finding of judge agreement on short-story rubrics demonstrate that the same approach scales to nonfiction editing, Sun's own domain; she imports the result by analogy. The evidence supports a narrower claim than "LLMs can learn taste": that an explicit, person-specific rubric measurably improves an LLM's editorial fit for that one person, while generative creative quality remains unaddressed by the same fix.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can readers reliably distinguish AI-written text from human writing? What limits language model accuracy in evaluating ideas? How can we detect and account for LLM involvement in academic writing?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 119 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Sun argues teaching an LLM a writer's own taste rubric turns it into a good editor even though it remains a bad creative writer