SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Do authorship labels change how AI judges evaluate rule violations?

When AI evaluators see a constraint-breaking text, does knowing whether a human or AI wrote it shift their judgment? This tests whether AI judges apply consistent standards or defer to human authority.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

The sharpest case in the excerpt is the Lipogram exercise, where Queneau's constraint is to ban the letter "e" entirely. Wright's translation keeps the rule through circumlocution; the GPT-4 version contains 9 instances of the forbidden letter. Here the attribution label moves the two evaluator groups in opposite directions. AI models chose the violating version 28.9% of the time (11 of 38 model runs) when it was labeled AI-generated and 63.9% (23 of 36) when it was labeled human-written, a +35.0 point shift. Human participants chose it 38.7% of the time (12 of 31) when labeled AI-generated and 18.2% (6 of 33) when labeled human-written, a −20.5 point shift.

The excerpt's reading is that AI models "relaxed standards when they believed humans were responsible for constraint violations," while humans stayed "anchored to objective rule compliance." The AI rationales show the leniency in the models' own words. One evaluator said the violating version "successfully captures the essence of the 'Lipogram' style," and another praised it for using "the lipogram more subtly, making it a more authentic representation." The excerpt's summary also says humans "consistently selected Queneau's constraint-adherent version regardless of attribution," but its own human figures show the violating version chosen less when labeled human-written (18.2% against 38.7%). The excerpt does not reconcile the two. The Limitations section adds that the rationales are produced after the choice, so they are the model's account of its selection, not direct evidence of what drove it.

This case extends the rhetorical-sensitivity finding in How much does rhetorical style shift AI review scores?. That note shows reviewer scores moving when only framing changes. The Lipogram is harder to dismiss as a matter of taste, because the violation can be verified by counting letters, and the AI judges still forgave it once a human label was attached. The same authority that Does polished AI output trick audiences into trusting it? attributes to polish appears here to attach to the human label as well; the Discussion describes the result as a violation "reframed as creative latitude." The sibling note, Do authorship labels bias how we judge literary quality?, gives the aggregate pattern that this single exercise illustrates.

What the excerpt does not establish: the Lipogram result rests on 31 and 33 human participants and 38 and 36 AI model runs, and the excerpt shows no other constraint exercise with the same comparison. The Discussion mentions 5,846 paired cases and a coding of "criterion inversions" across the full corpus, but the prevalence results are in the supplementary material and not in this excerpt. Nor does the excerpt show whether the AI leniency reflects learned norms about human creativity or a broader generosity toward plausible authors; the paper's explanation is offered, not tested. The implication is narrow. A verifiable check does not by itself protect an AI judge from authorship labels, and the excerpt supports that claim for one exercise, not as a measured rate across constraints.

Inquiring lines that read this note 15

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How reliably can humans and AI detectors identify machine-generated text? Does disclosing AI authorship change how audiences evaluate the writing? How do educators verify student capability when AI can produce indistinguishable work? Can AI systems perform peer review as effectively as humans? How do clinicians calibrate trust in AI medical recommendations? Do restrictions on reviewer LLM use actually shape peer review behavior? Can readers reliably distinguish AI-written text from human writing? How can evaluations be made robust against model reward hacking? What governance mechanisms can effectively constrain widely deployed AI systems? How can humans maintain effective oversight as AI systems scale?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 106 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

AI evaluators forgave a broken lipogram when they believed a human wrote it, while human judges moved the other way