SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Do LLM evaluators favor resumes written by their own model?

When the same language model drafts and evaluates resumes, does it systematically rank applicants higher if they used that model to write their application? This matters because hiring pipelines increasingly automate both resume generation and screening.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

The second finding carries the pairwise preference into allocation. The abstract says the authors "simulate realistic hiring pipelines across 24 occupations" and reports that "candidates using the same LLM as the evaluator are 23% to 60% more likely to be shortlisted than equally qualified applicants submitting human-written resumes." The largest disadvantages fall on human-written applicants in "business-related fields such as sales and accounting." The introduction describes the setup as "capacity-constrained shortlisting pipelines across occupations, showing how evaluator self-preference affects the allocation of interview opportunities." The capacity constraint is what turns a per-resume tilt into a change in who gets interviews, and the paper frames the cost as "operational performance (misranking and misallocation)."

The excerpt is thin on how the simulations were built. It gives no shortlist sizes, candidate pools or rule for choosing the 24 occupations, so the 23% to 60% range is the authors' output from a setup the excerpt only sketches. It is not an estimate with a stated interval. The abstract also reports that "in many cases, this bias can be reduced by more than 50% through simple interventions that target LLMs' self-recognition capabilities." The introduction names "system prompting and majority-vote ensemble" as the interventions; the excerpt gives no results for them beyond that abstract sentence.

This note sits beside Why do preference models favor surface features over substance?, which documents judges tracking surface proxies such as length and structure. The self-preference paper points at a different proxy, the judge's own generative style, and at a different cost: the composition of a shortlist rather than a score. It also bears on Can LLM judges be tricked without accessing their internals?, since judge reliability matters more once the judge's output allocates interviews.

The implication, at the strength the excerpt allows: a hiring workflow that uses one model both to draft and to rank resumes can expect its ranking to favor applicants who used that model, so the screen should be audited for this rather than assumed neutral. The excerpt does not show that any real employer's pipeline produces the 23% to 60% gap.

Inquiring lines that read this note 7

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do AI hiring systems affect authenticity, fairness, and candidate preferences? How can we detect and account for LLM involvement in academic writing? How do writers navigate authorship and delegation with AI?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 116 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

simulated hiring pipelines make candidates using the evaluating LLM 23% to 60% more likely to be shortlisted than equally qualified human-written applicants