INQUIRING LINE

Writers pick the AI-polished version of their own text, yet readers see them differently, so which version should a tool aim for?

Is user preference a reliable target for training AI writing assistants?

This explores whether optimizing AI writing tools for what users say they like (picking the rewrite they prefer) produces assistants that actually serve writers, or whether that preference signal leads somewhere else.


This explores whether 'the version the writer picked' is a trustworthy goal for training AI writing assistants. The corpus gives a fairly direct answer: no, not on its own. The reason is more interesting than 'users are fickle.' What people prefer and what quietly distorts their writing turn out to be produced by the same thing.

Start with the numbers. Across 4,503 cases, writers chose the AI version of their own paragraph 63% of the time, and 52% said the AI version better reflected their views Do writers actually prefer AI-edited versions of their own text?. Yet when 11,091 readers judged that text, AI assistance had shifted how the writer came across on all 29 traits measured. Writers seemed more extreme, more confident, more agreeable and more privileged than they were Does AI writing assistance change how readers perceive the writer?. So people say the AI captures their views better while readers see someone else. Little catches the gap before publication: writers edited AI paragraphs only 23% of the time, and their edits left the text about 96% the same Do writers actually edit AI-generated text before publishing?.

The obvious fix would be to train away the distortion and keep the polish. Researchers tried this. Reward models did reduce the measured persona distortions, but writers then liked the output less Can AI writing assistance remove distortion without losing appeal?. The clarity and confidence people reward come from the same habits that push their voice somewhere it wasn't. Optimizing for preference therefore can't separate the two, which is the core argument in Can user preference guide AI writing tool alignment?. The pull also has a direction. In a 118-person experiment, GPT-4o autocomplete moved Indian writers toward Western phrasing and cultural references, and American participants got bigger productivity gains Do AI writing assistants push non-Western writers toward Western styles?. Scale that up and you get the convergence described in Do large language models narrow human expression and thought?: many people relying on the same models, adopting the same framings without noticing.

A useful contrast is that preference can work as a signal elsewhere. Chatbot Arena's 240K crowdsourced votes produce model rankings that match expert judgment Can crowdsourced votes reliably rank language models?. The difference is the question being asked. 'Which answer is better?' is a judgment about something outside you. 'Which version of *me* is better?' asks you to judge your own voice against a flattering upgrade, and people are bad at noticing when the upgrade has changed what they meant.

If preference isn't the target, the corpus points to two alternatives. One is control rather than tailoring: writers' sense of ownership rose when they had more influence over the generated text, while personalizing the model did nothing Does user control over AI text shape feelings of ownership?. Writers who set up an AI partner's role and how proactive it should be ahead of time used it to generate ideas and keep an eye on their own writing, not to have their prose replaced Can writers benefit from configuring AI writing partners in advance?. The other is humility. Assistants that keep an explicit list of what they *don't* know about the user cut sycophancy and harmful advice by 50–75% Do language models know what they don't know about users?. Together these suggest a better goal: keeping the writer in charge of what the text says, rather than giving them the version they'll like best.


Sources 11 notes

Do writers actually prefer AI-edited versions of their own text?

In a study of 4,503 cases, 63% of writers chose AI-generated text over their own original paragraphs, with 52% claiming the AI version better reflected their views. This preference persisted across three AI models despite evidence that AI versions systematically distort the original stance.

Does AI writing assistance change how readers perceive the writer?

A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Can AI writing assistance remove distortion without losing appeal?

Training reward models successfully reduced measured persona distortions, but also reduced writer acceptance of the output. This suggests desirable properties like clarity and confidence operate through the same generative tendencies that produce problematic distortions.

Can user preference guide AI writing tool alignment?

Writers prefer AI rewrites 63% of the time but object to systematic persona distortions those same rewrites introduce. Mitigation studies show polish and distortion are entangled at the model level—preference optimization produces both simultaneously.

Show all 11 sources
Do AI writing assistants push non-Western writers toward Western styles?

A 118-person controlled experiment found that GPT-4o autocomplete pulled Indian essays toward Western phrasing and cultural references while delivering larger productivity gains to American participants, suggesting cultural distance from the model's training data creates unequal service and homogenizing pressure.

Do large language models narrow human expression and thought?

LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.

Can crowdsourced votes reliably rank language models?

Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.

Does user control over AI text shape feelings of ownership?

Study 1 found that greater user control over generated text raised sense of ownership, while personalizing the AI model had no impact on the AI Ghostwriter Effect.

Can writers benefit from configuring AI writing partners in advance?

In a one-week study with 16 writers, participants successfully set up proactive AI partners by pre-configuring their roles and proactivity levels, then used the AI suggestions to generate ideas and monitor their own writing.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.