How much does AI rewriting erase distinctive author voice?
Does heavy AI rewriting weaken the computational signals that identify individual authors? The question matters because it bears on whether AI assistance erases stylistic distinctiveness—a possible cost of polish and consistency.
Heavy rewriting by an AI writing assistant weakens the signals that let a computer tell one author from another, and the size of the loss depends on the register of the writing. The paper defines the Idiolect Erasure Rate (IER) as "the reduction in authorship-attribution accuracy following AI-assisted rewriting." Under heavy rewriting, the deep attributer LUAR loses 66.5 percentage points on personal blogs and 52.5 on Enron workplace email; the stylometric attributer loses 38.5 points on blogs, 28.7 on Enron and 10.0 on Reuters C50 news. The starting points are strong (LUAR reaches 0.815, 0.713 and 0.710 before rewriting), so the drop is measured from a working attributer. On news, where each journalist writes within a fixed beat, heavy rewriting "produces little deep erasure because topic remains highly predictive of authorship."
The authors read the loss as stylistic convergence, not a change of meaning. Three analyses point that way: rewritten texts become less distinguishable from one another, semantic similarity to the originals stays high, and function-word attribution, which carries little lexical content, also degrades. The loss rises with rewriting intensity. Grammar-only correction has much smaller effects, while even a prompt asking the assistant to preserve the author's voice removes most of the recoverable deep signal. Attribution also fails to accumulate: from original messages it approaches perfect accuracy within several messages, while from rewritten messages it saturates near 50 percent. The choice of attributer matters for a related reason. Shuffling word order drops LUAR from 0.710 to 0.205 but drops MiniLM only from 0.385 to 0.360, so the authors treat LUAR as style-sensitive and MiniLM as a topic-sensitive baseline.
This is the individual-level counterpart to the population result in Does AI writing make all writers sound the same?, which reports AI-assisted paragraphs rated more similar across writers than human-written ones. The paper says it asks a different question: "Rather than measuring homogenization across a population, we quantify the loss of attributable writing style for each author." It is also narrower than Does AI writing assistance change how readers perceive the writer?, which concerns perceived traits, where this paper asks whether the author can still be picked out at all. Its result also bears on Do users truly own the AI-generated content they produce?, since it measures a model's ability to attribute text, not anyone's sense of ownership. The detector claim the paper attaches to this result is taken up in Do rewrites that hide authorship also fool AI detectors?.
The excerpt does not establish what human readers recognize. The authors say IER "measures computational attributability, not human recognition," and list human-recognition studies as future work, so a colleague may still recognize a rewritten email. The setting is closed-set, which they call an upper bound on performance. The corpora are three pre-generative-AI sets of 20 to 50 authors, and the paper evaluates single-pass rewriting by three assistants. The primary rewriter is Qwen2.5-1.5B-Instruct; GPT-4o-mini and Gemini Flash are said to show "comparable levels" of deep erasure, but the excerpt gives no figures for them. Multilingual and longitudinal evaluations are listed as future work. The defensible claim is narrow: in these three registers, heavy rewriting by these assistants removes much of a model's ability to pick out an author's style. Whether that matters to a human reader is outside what was measured.
Inquiring lines that read this note 21
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can readers reliably distinguish AI-written text from human writing?- How much do writers actually edit AI text before publishing it?
- How much stylistic convergence is needed before a detector misses AI assistance?
- Do human readers still recognize authors after heavy AI rewriting?
- Why does topic structure protect authorship signals from AI erasure?
- Does AI assistance distort how readers perceive writer identity and demographics?
- Does polish in writing borrow authority that only expertise should carry?
- Does AI writing assistance distort a writer's authentic voice and persona?
- Does AI assistance distort how readers perceive a writer's voice?
- Does AI-generated writing feel polished while remaining harder to understand?
- Does the same rewriting that erases authorship also narrow measurable AI text markers?
- Can detection systems identify AI text rewritten to match human author style?
- Can AI-rewritten text still be detected as machine-modified?
- Can AI detectors confuse distinctive writing style for machine authorship?
- Does AI writing assistance make different authors sound more alike?
- Does asking AI to preserve voice recover lost authorship signals?
- Can proofreading tools preserve writer voice better than full rewriting features?
- What mechanisms explain why rivalry reduces certain writing tasks for some writers?
- When AI becomes invisible in writing tools, do writers stop disclosing it?
- How does reliance on AI change when writers own the final product?
- How does heavy AI use change the thinking that happens during writing?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does AI writing make all writers sound the same?
When writers use AI assistance, do their distinct voices converge toward a generic style? This matters because readers rely on voice to identify and distinguish among individual writers.
population-level, rater-based homogenization; this paper measures each author's attributable signal instead
-
Does AI writing assistance change how readers perceive the writer?
Explores whether AI-assisted writing systematically alters reader impressions of the writer's political views, competence, emotion, and demographic identity. Understanding this matters because perception shapes trust and influence in public discourse.
perceived traits versus whether the author is still attributable at all
-
Do users truly own the AI-generated content they produce?
When people use AI to create outputs, do they experience genuine authorship and ownership of what's produced, or does the continuous interaction loop create a gap between what they feel and what they claim?
IER measures computational attributability, not felt or claimed ownership
-
Do rewrites that hide authorship also fool AI detectors?
A paper claims heavily rewritten AI text becomes both unattributable to humans and undetectable as AI-assisted, but tests only the attribution half. Does rewriting actually evade detection systems?
the detector claim built on this attribution result, untested in the excerpt
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Assistant Erased You: Measuring Loss of Authorship Signals in AI-Mediated Communication
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- "It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- StoryScope: Investigating idiosyncrasies in AI fiction
Original note title
Heavy AI rewriting weakens authorship signals far more in blogs and email than in topic-structured news — the Idiolect Erasure Rate measures the drop