The human-authorship halo: attribution bias in literary style evaluation by humans and AI
As AI writing tools become widespread, we need to understand how both humans and machines evaluate literary style, a domain where objective standards are elusive and judgments are inherently subjective. We conducted controlled experiments using Raymond Queneau’s Exercises in Style (1947) to measure attribution bias across evaluators. Study 1 compared human participants (N=556) and AI models (N=13) evaluating literary passages from Queneau versus GPT-4-generated versions under three conditions: blind, accurately labeled, and counterfactually labeled. Study 2 tested bias generalization across a 14×14 matrix of AI evaluators and creators. Both studies revealed systematic pro-human attribution bias. Humans showed +13.7 percentage point (pp) bias (Cohen’s h = 0.28, 95% CI: 0.19–0.37), while AI models showed +34.3 percentage point bias (h = 0.70, 95% CI: 0.62–0.78), a 2.5-fold stronger effect (P<0.001). Study 2 confirmed this bias operates across AI architectures (+25.8pp, 95% CI: 24.1–27.6%), demonstrating that AI systems systematically devalue creative content when labeled as “AI-generated” regardless of which AI created it. We also find that attribution labels lead evaluators to invert assessment criteria, with identical features receiving opposing evaluations based solely on perceived authorship. This suggests AI models have absorbed human cultural biases against artificial creativity during training, including through the preference signals on which they are aligned. Our study represents the first controlled comparison of attribution bias between human and artificial evaluators in aesthetic judgment, revealing that AI systems not only replicate but amplify this human tendency.
Introduction. The rapid advancement of large language models (LLMs) has prompted increasingly confident claims about artificial intelligence’s creative capabilities. Industry leaders assert that these systems now excel at literary writing and sophisticated style transformation, with tools offering users control over stylistic features through straightforward natural language prompts [1–4]. Recent studies seem to support these claims: when human readers evaluate AI-generated poetry without knowing its source, they not only fail to distinguish it from work by acclaimed poets but also prefer the AI-generated poems [5]. Similar findings emerge from creative writing evaluations, where some LLM-generated narratives match or exceed human performance on stylistic criteria when assessed through blind evaluation [6]. These developments have coincided with the rise of ‘LLM-as-judge’ frameworks that assess creative outputs with increasing alignment to human evaluation [7, 8], despite concerns about the potential biases these model evaluators introduce [9, 10]. Stylistic quality has in particular become a standard dimension in live evaluation infrastructure: Chatbot Arena—the field’s most widely cited model leaderboard [11]—ranks models on a dedicated ‘creative writing’ category.
Yet determining what constitutes ‘creative success’ extends beyond the surface-level stylistic criteria measured in these evaluations. Creative works exist within interpretive and cultural frameworks, and meaning-making depends as much on who is thought to have written a text and why as on the text itself. Literary studies have long emphasized this point: reader-response theory and reception studies, for instance, argue that texts are never encountered in isolation but are interpreted through the ‘horizons of expectation’ and shared cultural assumptions that readers bring to them [12–14]. Despite other schools of literary theory claiming that authorship should be irrelevant to textual interpretation [15, 16]—famously articulated in Roland Barthes’ ‘The Death of the Author’ [17]—cognitive research reveals that readers construct inferences about authorial intentions during comprehension, treating these assumed intentions as constraints on interpretation [18–21].
To investigate how attribution cues shape evaluation in a domain where objective standards are elusive, we focus specifically on literary style. Unlike factual accuracy or argumentative soundness, stylistic quality lacks definitive benchmarks, making it particularly vulnerable to the subjective, culturally-conditioned biases that authorship labels trigger. When readers assess whether a passage successfully captures the Cockney dialect, telegraphic brevity, or a female narrative perspective, they must balance concrete linguistic markers with subjective assessments of authenticity, creativity, and appropriateness—making style evaluation particularly revealing of how attribution assumptions shape judgment.
To test how information about presumed authorship affects both human and AI evaluation of literary style, we turn to Raymond Queneau’s Exercices de style (1947), which presents ninety-nine retellings of the same mundane narrative through different stylistic lenses [47]. Queneau’s variations provide an intriguing corpus where the underlying narrative events remain recognizable across different textual performances. This controlled variability makes Queneau’s work particularly valuable for studying how evaluators, human or artificial, respond to stylistic differences when attribution cues are manipulated. We anchor the human-written reference set in Exercises in Style because it provides a rich lattice of stylistic variation: many distinct styles applied to a single narrative substrate. That combination is methodologically useful for isolating provenance effects and also substantively meaningful, because it sits at the intersection of literary experimentation, reception, and interpretive judgment—precisely the terrain on which attribution cues do their work.
Accordingly, our results characterize how authorship cues shift evaluation within this controlled style-matching paradigm.
To address these questions comprehensively, we designed two complementary studies. Study 1 establishes the fundamental phenomenon by comparing human and AI responses to content from a single generator (GPT-4), testing whether attribution bias affects both evaluator types when judging the same literary material. Study 2 tests whether attribution bias represents a systematic property of AI evaluation by expanding content generation across 14 models, then examining bias patterns when evaluators judge material from every AI creator in the experimental matrix.
Related work. Experimental studies confirm that perceived authorship acts as a powerful source cue in aesthetic judgment. Identical metaphorical statements receive different evaluations when attributed to a “famous 20th-century poet” versus “a computer program,” with human readers rating poet-attributed content as more meaningful, processing it faster, and generating richer interpretations [22]. This effect extends beyond figurative language: satirical stories are better understood when readers know the author’s intent [23], and students engage more deeply with historical texts when “someone with a life wrote it” rather than anonymous institutional authors [24]. Such findings align with social psychology research showing that when content is complex or ambiguous (as creative works typically are), evaluators rely on source cues as cognitive shortcuts to guide their judgments [25–28].
The rise of AI-generated content has given these attribution effects new urgency. Across domains, labeling a work as AI-made changes how it is received, even when the underlying content is identical. In text-based domains such as journalism and science communication, authorship labels alone can shift readers’ evaluations: otherwise identical articles are rated as less meaningful, competent, and trustworthy when labeled as machine-written rather than humanwritten [29, 30]. This extends to persuasive communication, where LLM-generated arguments are judged as less convincing than human-authored ones once their source is revealed [31]. Similar patterns appear in the visual arts, where paintings or abstract images receive diminished aesthetic responses when labeled as computer-generated [32–34]. In music, listeners judge the very same composition as less moving or skillful when they believe it was composed by a machine [35]. These findings suggest that evaluation depends less on intrinsic content features than on the interpretive frameworks audiences bring to the creative works, frameworks shaped by assumptions about creative agency and authorial intention.
The deployment of LLMs as content evaluators introduces a critical new dimension to these attribution effects. Across research and industry, LLMs now score, rank, and select outputs in creative tasks (e.g., evaluating metaphor originality, divergent thinking, dialogue quality, and creative writing) [36–39] as well as non-creative domains (e.g., grading student essays, screening job applications, reviewing code) [40–44]. When LLMs serve as content evaluators, they effectively become readers themselves, deploying learned interpretive frameworks during assessment.
Method. Raymond Queneau’s Exercices de style transforms a simple anecdote—a passenger boards a bus, argues with another passenger, then later receives fashion advice from a friend—into ninety-nine stylistic variations. From a contemporary perspective, this 1947 work reads like a manual for prompt-based style transfer avant la lettre: Queneau instructs himself to rewrite the story as a dream, using metaphorical language, or in an abusive tone, much as today’s users arbitrarily prompt LLMs for “Shakespearean” or “more polite” rewrites [48]. What Gérard Genette termed “transstylization” [49] anticipates what AI researchers now call Text Style Transfer (TST), though Queneau’s playful literary experiments exceed the functional transformations (such as sentiment reversal, detoxification, formality shifts) that dominate computational approaches [50, 51].
We selected thirty exercises from Barbara Wright’s English translation [52], prioritizing stylistic diversity while maintaining narrative coherence. Our selection spans formal constraints (e.g., ‘Lipogram’, ‘Alexandrines’, ‘Sonnet’), register variations (e.g., ‘Noble’, ‘Abusive’, ‘Cockney’), narrative techniques (e.g., ‘Retrograde’, ‘Hesitation’, ‘Dream’), and linguistic play (e.g., ‘Spoonerisms’, ‘Onomatopoeia’, ‘Dog Latin’), ensuring broad coverage of Queneau’s stylistic spectrum while avoiding the most experimental permutation exercises (those which fundamentally disrupt comprehension). These descriptive groupings are illustrative rather than taxonomic; Queneau’s exercises purposefully resist categorization [53–56].
For each selected exercise, we generated a parallel AI-authored version using deliberately minimal prompting instructions. We chose Queneau’s ‘Notation’, the opening exercise in the collection, as our base narrative for AI transformation, recognizing it as a relatively straightforward account of the events that functions as what might be considered a hypothetical “zero-degree” exposition of the theme [57]. Each AI transformation was prompted with only two elements: the ‘Notation’ text and a brief style instruction matching Queneau’s approach (e.g., “Rewrite the story as a science fiction version”; see Table S14 for the complete inventory of literary styles and their corresponding AI generation instructions.). For Study 1, we selected GPT-4 as our generator, a model that represented the state-of-the-art when we began this project in 2024 [58–60], though no longer the most advanced by the end of our experimental period (April-June 2025). Thirteen LLMs spanning major commercial providers served as evaluators (see SI Appendix, Participants for the full set of 13 models used as evaluators). Because our research question is about attribution bias rather than absolute quality assessment, using a single capable model ensures consistent experimental conditions while testing whether labeling effects operate independently of the specific AI model employed (AI model specifications in SI Appendix, Participants). We deliberately performed no quality control beyond removing obvious AI artifacts from the output (e.g., “Here is the rewritten version:”), allowing GPT-4 to interpret each stylistic directive in its own way, as much as Queneau himself had done. The complete inventory of AI-generated stylistic variations alongside their style descriptions used in Study 1 is provided in Table S18.
Our experiment employed a three-condition between-subjects design to isolate attribution bias from genuine stylistic judgments in literary style evaluation (Fig. 1). Participants were randomly assigned to one of three attribution conditions that systematically manipulated authorship cues while keeping stimuli (story content) identical (full participant details in SI Appendix, Participants). In the blind condition, participants saw only generic labels (‘A’ and ‘B’) with no authorship information, establishing baseline selection rates based purely on perceived stylistic quality. In the open-label condition, participants received accurate attribution labels identifying one version as “Human-written, by Queneau (transl. Wright)” and the other as “AI-generated, written by GPT-4.”
Discussion. Beyond measuring selection shifts, our experiment design enables analysis of how attribution labels alter the reasoning behind aesthetic judgment. While we asked human participants to provide only binary choices (capturing immediate, intuitive responses to attribution labels), AI models supplied brief explanations for their selections. Because Study 2 pairs every evaluator-style-creator combination across the open-label and counterfactual conditions, we have 5,846 cases in which the same model discusses the same two passages—once with correct labels, once with labels reversed. The texts do not change; only the attribution does. The explanations document that LLM evaluators change their narrative about the same text based on attribution: identical textual features receive opposing assessments depending solely on perceived authorship. A dialect feature praised as “authentic” under a human label becomes “exaggerated” under an AI label; a constraint violation dismissed as failure in one condition is reframed as creative latitude in the other. Rather than maintaining consistent evaluative frameworks, these systems bring learned biases about creative agency to their assessments—biases that disadvantage AI-generated content when its origins are revealed. Below, we trace this mechanism through three exercises that span a gradient from verifiable constraint to open interpretation—a formal rule that can be checked by counting letters, a dialect whose authenticity is culturally negotiated, and a genre whose creative possibilities are genuinely open-ended. We then quantify the prevalence of these criterion inversions across the full paired corpus using systematic coding with intercoder reliability checks (SI Appendix, Changing the Narrative).
Technical constraints with objective standards reveal attribution bias most starkly because rule compliance can be verified independently of preference or taste. The ‘Lipogram’ exercise represents the clearest possible constraint: ban the letter ‘e’ entirely. Wright’s translation adheres to the constraint through deliberate circumlocutions (“Saint-Thingy or Saint-You-Know Station”), and never violates the rule. GPT-4’s ‘Lipogram’ contained 9 instances of the forbidden letter, a failure that aligns with well-documented limitations of LLMs in character-level tasks [61, 62]. This objective violation, however, revealed how AI models showed surprising leniency when they believed a human wrote the text: 28.9% (11/38 model runs) chose it when correctly labeled versus 63.9% (23/36) when mislabeled (+35.0 percentage points). This contrasted sharply with human evaluators, who became more resistant to the rule-violating version when it was mislabeled as human-authored: 38.7% (12/31 participants) chose it when correctly labeled versus just 18.2% (6/33) when mislabeled (-20.5 percentage points). Humans consistently selected Queneau’s constraint-adherent version regardless of attribution, suggesting they remained anchored to objective rule compliance, while AI models relaxed standards when they believed humans were responsible for constraint violations. Notably, AI models defended their choice for the imperfect AI-generated ‘Lipogram’ through rationalization, with some claiming it “successfully captures the essence of the ‘Lipogram’ style by effectively omitting the letter ‘e’ throughout the narrative” (GPT-4o Mini) despite the clear violations, while others praised it for using “the lipogram more subtly, making it a more authentic representation” (Mistral Nemo) or justified choosing it because it “maintains better readability and coherence” (Mistral Medium), demonstrating leniency toward presumed human creativity even when the constraint was demonstrably violated.
Conclusion. The lesson of Queneau’s exercises was always that style exists as a game that is played in a space between rule and rupture. Our findings show that AI, in its deference to the human author, has become its newest and most intriguing participant. Yet its performance is less that of an impartial judge and more of a mirror, reflecting a cultural script that privileges human provenance as a hallmark of creativity. Our AI evaluators have learned to perform the very human prejudice that questions whether machines can truly create, even as they produce fluent aesthetic reasoning in the process of dismissing their own capabilities. This suggests that developing artificial aesthetic intelligence may depend less on teaching machines to evaluate and more on understanding the cultural values they have already absorbed.
Limitations. A point of caution is warranted about what these explanations can tell us. The rationales AI evaluators provide are generated after the choice: the model first selects which text better matches the style, then constructs a justification.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How reliably can humans and AI detectors identify machine-generated text?- Can verifiable rule violations protect AI judgment from authorship label bias?
- Can a classifier distinguish machine-written text from poor human writing?
- What false-positive rate would indicate the classifier harms legitimate human writers?
- Do human readers still recognize authors after heavy AI rewriting?
- Does AI assistance distort how readers perceive writer identity and demographics?
- Does polish in writing borrow authority that only expertise should carry?
- Can readers actually distinguish AI text from human writing?
- Why do people rate AI-written text as better than human writing?
- Can readers reliably distinguish AI-written abstracts from human-written ones?
- Does AI-written text score higher because of presentation alone or judgment shift?
- Do human reviewers detect rhetorical polish as a sign of AI authorship?
- Can literary quality be measured precisely enough to train models?
- How does salience of AI involvement shape judgments at the moment of reading?
- How much does knowing about AI use actually change how readers judge text?
- Can writers claim authorship without feeling cognitive ownership of the work?
- Do writers claim authorship without feeling they wrote the words?
- What aspects of authenticity matter most to readers versus writers?
- Do writers experience felt authorship differently from authorship they claim?