SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Do LLM raters show hidden demographic preferences that disclosure erases?

Explores whether language models systematically favor certain demographic groups when their AI involvement is not disclosed, and whether that preference disappears under transparency. This matters because it reveals potential fragility in AI alignment training.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

Only the LLM raters in this study show demographic interaction effects. Human disclosure penalties were "relatively uniform across authors of different races and genders." The two LLMs were not uniform. GPT-4o-mini "showed a pronounced preference for Black authors in the control condition, which diminished when AI involvement was disclosed." Qwen2.5-7B-Instruct "similarly favored woman authors in the absence of disclosure, a gender bias that disappeared when AI assistance was acknowledged." The authors note that both favored groups "are historically marginalized." They name the pattern "vanishing alignment," described as a case "where the social and ethical calibration of model behavior becomes fragile under changing contextual cues."

The authors' explanation is hedged. The preferences "may reflect an alignment-driven over-correction, in which models trained with human feedback disproportionately reward underrepresented identities." They tie the pattern to Hofmann et al., who found that preference-aligned LLMs projected "overtly positive stereotypes" toward African American English speakers "while maintaining covertly negative biases that surfaced in indirect prompts." For why the effect appears in LLMs and not in humans, they offer two conjectures. LLM raters "might produce more stable, patterned outputs across multiple runs, making their biases more legible." And news-genre expectations of objectivity may amplify penalties for AI involvement, "reducing the salience of the author's race or gender." Neither conjecture is tested in the excerpt.

Set against the nearest notes, this note moves the question from the text to the judge. Does AI writing assistance change how readers perceive the writer? reports that AI assistance distorts demographic markers in what readers see. Here the concern is that the scoring models' own demographic preferences depend on a contextual cue. The cue is the same kind of lever as in Does telling people an AI wrote something actually stop them from believing it?, which found disclosure raises scrutiny without collapsing persuasion. In this study the same cue switches off a demographic preference in two LLMs. The human sample does not share that pattern, since its disclosure penalty is uniform across groups.

The excerpt does not establish the size of the demographic effects. It says the authors fit "linear models with interaction terms" but reports no estimates, test statistics, or number of LLM ratings per condition. The abstract refers to "LLM raters" generally, while the discussion names only two models, so the excerpt does not show that they are the full set tested. The over-correction account is a possibility the authors raise, not a finding. The implication is conditional. If the pattern holds beyond these two models, an LLM used to score writing would need audits that vary contextual cues such as disclosure, not only the author's identity. On this excerpt alone, the claim is that two models showed the pattern under one news article and one disclosure wording.

Inquiring lines that read this note 25

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do AI hiring systems affect authenticity, fairness, and candidate preferences? Why do confident AI outputs mislead human trust calibration? Can base models hide emergent misalignment through alignment training? How does AI-generated content create social proof without authentic interaction? Does disclosing AI authorship change how audiences evaluate the writing? How can we detect and account for LLM involvement in academic writing? Do language models reason through disagreement or only accommodate it? How do writers navigate authorship and delegation with AI? How reliably can humans and AI detectors identify machine-generated text? Can persona profiles improve LLM prediction accuracy and consistency? How does AI adoption reshape collaboration patterns in knowledge work?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 93 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM raters favor Black or women authors when AI use goes undisclosed, and the preference vanishes once it is revealed