INQUIRING LINE

Tell people an AI helped write something and they score it lower, even when the words are identical.

Can disclosure of AI involvement change how evaluators score writing quality?

This explores whether telling evaluators that AI was involved in a piece of writing changes the scores they give, even when the text itself is the same.


This explores whether a label saying "AI helped write this" changes how people, and AI judges, score the same text. The corpus says yes. How much depends on who is judging, what kind of writing it is, and what the judge thinks AI use says about the writer. The cleanest test used an identical news article shown with and without an AI disclosure statement. Both human raters and LLM raters scored the disclosed version lower Does disclosing AI assistance make readers trust articles less?. The penalty was real and consistent, but small: less than 0.15 points on a 7-point scale. So for informational writing, disclosure nudges scores down rather than tanking them. It's also notable that the LLM raters showed the same bias. AI judges pick up human-like reactions to an AI label.

The penalty gets much larger when the writing is supposed to come from a person who cares. When readers learned AI wrote a text, they rated it lower on trust, caring and likability, and the steepest drops came in interpersonal writing How does revealing AI authorship change reader trust?. Readers treated AI use there as breaking a social expectation, because they don't believe AI can mean the empathy it expresses. The reader also matters. People with higher AI literacy showed smaller drops after disclosure, and some even viewed AI use positively Does AI literacy reduce the damage from AI disclosure?. These attitudes also explain the tension over who should disclose: readers consistently judge disclosure more necessary than writers do Do readers and writers differ on AI disclosure necessity?.

The effect doesn't need a formal label. Suspicion alone does the work. Admissions officers could often spot AI-written essays and rated the essays they believed were AI lower Do admissions officers penalize essays they suspect are AI-written?. In a real admissions pool, applicants whose essays were flagged as likely AI were admitted at lower rates than comparable applicants, even though AI had improved their essays Does AI essay use hurt admissions chances despite quality gains?. The authors propose this suspicion as the likely explanation rather than proving it directly. Still, it points to a striking result: better writing can lose if the evaluator thinks a machine produced it.

The most surprising finding is that human and AI judges react to authorship labels in opposite directions. When a lipogram (a text written without using a particular letter) broke its own rule, AI judges picked it 35 percentage points more often when told a human wrote it. Human judges picked it 20 points less often Do authorship labels change how AI judges evaluate rule violations?. AI judges seem to go easy on human work, while humans hold one another to the letter of the rules. This matters as LLMs increasingly grade essays and review papers, because the same disclosure can push scores in opposite directions depending on who grades.

One complication is that disclosure isn't purely a bias acting on unchanged text. AI assistance really does change the writing. It shifted readers' impressions of writers on all 29 dimensions tested, making writers seem more extreme, more confident and more privileged Does AI writing assistance change how readers perceive the writer?. Writers edited AI suggestions only 23% of the time, and lightly when they did Do writers actually edit AI-generated text before publishing?. So when evaluators mark down AI-assisted work, some of that is a reaction to the label and some is a reaction to real changes in voice. The corpus separates the two only in the identical-text study, where the label effect turns out to be small.


Sources 9 notes

Does disclosing AI assistance make readers trust articles less?

Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.

How does revealing AI authorship change reader trust?

A study of 261 readers found that disclosing AI authorship consistently lowered perceived trustworthiness, caring, and likability, with the steepest drops in interpersonal writing like personal interaction. Readers saw AI as incapable of genuine empathy, viewing its use as a violation of social expectations.

Does AI literacy reduce the damage from AI disclosure?

In a 261-person study, readers with higher self-reported AI literacy showed smaller negative shifts in perception after learning AI was used, and some expressed positive attitudes toward AI use. Literacy appears to act as a boundary condition on the broader disclosure penalty.

Do readers and writers differ on AI disclosure necessity?

A 727-person vignette study found readers consistently rated AI disclosure as more necessary than writers did. Disclosure seemed most necessary when AI text was directly incorporated and irreplaceable, while writer effort had no effect on these judgments.

Do admissions officers penalize essays they suspect are AI-written?

In an experiment, admissions officers could often discriminate AI from human essays and rated essays they believed to be AI-generated lower than those believed human-written. The authors frame this as a plausible explanation for the observed admissions penalty, though the link remains proposed rather than directly measured.

Show all 9 sources
Does AI essay use hurt admissions chances despite quality gains?

Among 7,500 applications to a public policy master's program, majority of 2025 applicants submitted AI-generated essays despite explicit prohibition. These applicants were admitted at lower rates than similar applicants without detected AI use, despite AI improving essay quality.

Do authorship labels change how AI judges evaluate rule violations?

AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.

Does AI writing assistance change how readers perceive the writer?

A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.