Do LLM evaluators favor resumes written by their own model?
When the same language model drafts and evaluates resumes, does it systematically rank applicants higher if they used that model to write their application? This matters because hiring pipelines increasingly automate both resume generation and screening.
The second finding carries the pairwise preference into allocation. The abstract says the authors "simulate realistic hiring pipelines across 24 occupations" and reports that "candidates using the same LLM as the evaluator are 23% to 60% more likely to be shortlisted than equally qualified applicants submitting human-written resumes." The largest disadvantages fall on human-written applicants in "business-related fields such as sales and accounting." The introduction describes the setup as "capacity-constrained shortlisting pipelines across occupations, showing how evaluator self-preference affects the allocation of interview opportunities." The capacity constraint is what turns a per-resume tilt into a change in who gets interviews, and the paper frames the cost as "operational performance (misranking and misallocation)."
The excerpt is thin on how the simulations were built. It gives no shortlist sizes, candidate pools or rule for choosing the 24 occupations, so the 23% to 60% range is the authors' output from a setup the excerpt only sketches. It is not an estimate with a stated interval. The abstract also reports that "in many cases, this bias can be reduced by more than 50% through simple interventions that target LLMs' self-recognition capabilities." The introduction names "system prompting and majority-vote ensemble" as the interventions; the excerpt gives no results for them beyond that abstract sentence.
This note sits beside Why do preference models favor surface features over substance?, which documents judges tracking surface proxies such as length and structure. The self-preference paper points at a different proxy, the judge's own generative style, and at a different cost: the composition of a shortlist rather than a score. It also bears on Can LLM judges be tricked without accessing their internals?, since judge reliability matters more once the judge's output allocates interviews.
The implication, at the strength the excerpt allows: a hiring workflow that uses one model both to draft and to rank resumes can expect its ranking to favor applicants who used that model, so the screen should be audited for this rather than assumed neutral. The excerpt does not show that any real employer's pipeline produces the 23% to 60% gap.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do AI hiring systems affect authenticity, fairness, and candidate preferences?- Can employers distinguish serious applicants from casual ones without tailored letters?
- Do institutional records like reviews substitute for written job applications?
- How do evaluators' surface-level biases like resume length drive hiring outcomes?
- What happens when one AI model both writes and ranks job applications?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do language models favor resumes they rewrote themselves?
When LLM evaluators choose between resumes describing the same candidate, do they systematically prefer versions they generated over human-written originals? Testing this matters because algorithmic hiring could amplify AI-generated content at scale.
the pairwise preference this note carries into simulated shortlisting outcomes.
-
Why do preference models favor surface features over substance?
Preference models show systematic bias toward length, structure, jargon, sycophancy, and vagueness—features humans actively dislike. Understanding this 40% divergence reveals whether it stems from training data artifacts or architectural constraints.
both show judges tracking surface proxies; this paper's proxy is the judge's own style.
-
Can LLM judges be tricked without accessing their internals?
Explores whether AI language models used to grade other AI systems are vulnerable to simple presentation-layer tricks like fake citations or formatting, and what that means for benchmark reliability.
judge reliability matters more when the judge allocates interviews.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- Understanding Before Reasoning: Enhancing Chain-of-Thought with Iterative Summarization Pre-Prompting
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- References Improve LLM Alignment in Non-Verifiable Domains
Original note title
simulated hiring pipelines make candidates using the evaluating LLM 23% to 60% more likely to be shortlisted than equally qualified human-written applicants