SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Do language models favor resumes they rewrote themselves?

When LLM evaluators choose between resumes describing the same candidate, do they systematically prefer versions they generated over human-written originals? Testing this matters because algorithmic hiring could amplify AI-generated content at scale.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

The central claim is that an LLM asked to choose between resumes favors the one it generated itself, even when both describe the same candidate and content quality is held constant. The paper calls this "AI self-preference bias" and reports it from a "large-scale controlled resume correspondence experiment" over 2,245 human-written resumes collected "prior to the widespread adoption of generative AI." The abstract gives "self-preference bias ranging from 67% to 82% across major commercial and open-source models" and says "the bias against human-written resumes is particularly substantial." The results section reports the comparison on another scale: eight of nine LLMs show LLM-vs-Human self-preference, "with magnitudes ranging from 26% to 98%."

The mechanism the paper gives is stylistic and endogenous. Self-preference "emerges endogenously from AI-AI interactions, in which the model's own evaluative behavior systematically favors outputs aligned with its generative patterns." The authors treat the evaluator as a binary classifier and test it against statistical parity (whether selection rates differ by source of generation) and equal opportunity (whether they differ conditional on merit). The results section says the strength of self-preference "increases with model size," which "may indicate that larger models are more sensitive to stylistic features resembling their own outputs." That is offered as a possible reading, not a tested cause.

This is a different bias from the one in Can LLM judges be fooled by fake credentials and formatting?, where authority and beauty cues are added to text by someone gaming the judge. Self-preference needs no added content; it comes from the match between judge and candidate. The excerpt says it is "not addressed by existing safeguards focused on demographic disparities," so the benchmark-gaming framing in Can LLM judges be tricked without accessing their internals? covers only part of the risk. The closest parallel is Can user preference guide AI writing tool alignment?: both describe a pull toward model-styled text, here measured on evaluators rather than writers.

The excerpt does not show how any employer's live screening would behave; the evidence is paired counterfactual resumes judged by LLMs, with no human-recruiter comparison. The LLM-vs-LLM results are "considerably more heterogeneous across models," but their figures fall outside the excerpt. It also does not explain how the abstract's 67% to 82% range relates to the 26% to 98% range in the results. The defensible reading is narrower than the title: under controlled pairing, the evaluators tested lean toward their own output, and whether that carries into human-supervised hiring remains open.

Inquiring lines that read this note 24

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do AI hiring systems affect authenticity, fairness, and candidate preferences? How can we detect and account for LLM involvement in academic writing? How can AI systems reliably guide voters without introducing political bias? What gaps exist between benchmark performance and real deployment outcomes? How do writers navigate authorship and delegation with AI?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 119 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llm evaluators prefer resumes they generated themselves when content quality is controlled — self-preference in algorithmic hiring