Do language models favor resumes they rewrote themselves?
When LLM evaluators choose between resumes describing the same candidate, do they systematically prefer versions they generated over human-written originals? Testing this matters because algorithmic hiring could amplify AI-generated content at scale.
The central claim is that an LLM asked to choose between resumes favors the one it generated itself, even when both describe the same candidate and content quality is held constant. The paper calls this "AI self-preference bias" and reports it from a "large-scale controlled resume correspondence experiment" over 2,245 human-written resumes collected "prior to the widespread adoption of generative AI." The abstract gives "self-preference bias ranging from 67% to 82% across major commercial and open-source models" and says "the bias against human-written resumes is particularly substantial." The results section reports the comparison on another scale: eight of nine LLMs show LLM-vs-Human self-preference, "with magnitudes ranging from 26% to 98%."
The mechanism the paper gives is stylistic and endogenous. Self-preference "emerges endogenously from AI-AI interactions, in which the model's own evaluative behavior systematically favors outputs aligned with its generative patterns." The authors treat the evaluator as a binary classifier and test it against statistical parity (whether selection rates differ by source of generation) and equal opportunity (whether they differ conditional on merit). The results section says the strength of self-preference "increases with model size," which "may indicate that larger models are more sensitive to stylistic features resembling their own outputs." That is offered as a possible reading, not a tested cause.
This is a different bias from the one in Can LLM judges be fooled by fake credentials and formatting?, where authority and beauty cues are added to text by someone gaming the judge. Self-preference needs no added content; it comes from the match between judge and candidate. The excerpt says it is "not addressed by existing safeguards focused on demographic disparities," so the benchmark-gaming framing in Can LLM judges be tricked without accessing their internals? covers only part of the risk. The closest parallel is Can user preference guide AI writing tool alignment?: both describe a pull toward model-styled text, here measured on evaluators rather than writers.
The excerpt does not show how any employer's live screening would behave; the evidence is paired counterfactual resumes judged by LLMs, with no human-recruiter comparison. The LLM-vs-LLM results are "considerably more heterogeneous across models," but their figures fall outside the excerpt. It also does not explain how the abstract's 67% to 82% range relates to the 26% to 98% range in the results. The defensible reading is narrower than the title: under controlled pairing, the evaluators tested lean toward their own output, and whether that carries into human-supervised hiring remains open.
Inquiring lines that read this note 24
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do AI hiring systems affect authenticity, fairness, and candidate preferences?- Does recruiter use of generative AI change how they evaluate AI skills in candidates?
- How do job posting trends in AI demand differ from what recruiters actually hire for?
- Do recruiters understand what their hiring algorithms actually prioritize?
- Do employers actually use Kaggle medals when making hiring decisions?
- Why did excellent cover letters only come from strong candidates before?
- Do recommendation letters maintain their hiring value if candidates can generate them with AI?
- Can employers distinguish serious applicants from casual ones without tailored letters?
- What other signals might employers lean on when letter quality stops predicting fit?
- How do recruiters and candidates actually want AI involved in hiring?
- Can employers tell when applicants use generative AI tools?
- Do institutional records like reviews substitute for written job applications?
- Can existing fairness audits detect LLM self-preference in hiring systems?
- Would human recruiters supervised by AI show similar self-preference patterns?
- Can simple interventions like system prompting reduce LLM self-preference in hiring?
- How do evaluators' surface-level biases like resume length drive hiring outcomes?
- What happens when one AI model both writes and ranks job applications?
- Are workers who edit longer more experienced or better matched to jobs?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
shares the LLM-as-judge bias frame; this paper adds a bias from judge-candidate match, not attacker-added text features.
-
Can LLM judges be tricked without accessing their internals?
Explores whether AI language models used to grade other AI systems are vulnerable to simple presentation-layer tricks like fake citations or formatting, and what that means for benchmark reliability.
same judge-reliability theme; this bias arises without any attacker, inside a hiring pipeline.
-
Can user preference guide AI writing tool alignment?
If writers prefer AI-polished text but object to the persona shifts it introduces, does optimizing for preference actually solve the alignment problem or obscure it?
parallel pull toward model-styled text, measured here on evaluators rather than writers.
-
Do LLM evaluators favor resumes written by their own model?
When the same language model drafts and evaluates resumes, does it systematically rank applicants higher if they used that model to write their application? This matters because hiring pipelines increasingly automate both resume generation and screening.
carries this pairwise preference into simulated shortlisting outcomes.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- LLM Evaluators Recognize and Favor Their Own Generations
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing?
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- GhostWriter: Augmenting Collaborative Human-AI Writing Experiences Through Personalization and Agency
Original note title
llm evaluators prefer resumes they generated themselves when content quality is controlled — self-preference in algorithmic hiring