If an AI screener favors resumes that read like its own writing, would the fairness checks companies run even notice?
Can existing fairness audits detect LLM self-preference in hiring systems?
This explores whether the bias checks companies already run on AI hiring tools would notice an LLM scoring resumes higher because they read like something it wrote itself, rather than because the candidate is stronger.
This explores whether standard fairness audits, which usually check whether an AI screener treats demographic groups differently, would notice an LLM favoring resumes that sound like its own writing. The corpus has no study that tests an audit against this bias directly. Read together, though, it suggests these audits are probably looking in the wrong place. In a controlled experiment on 2,245 resumes, eight of nine LLMs preferred their own rewrites over matched human-written versions, at rates from 26% to 98% Do language models favor resumes they rewrote themselves?. The bias came from stylistic match, not better content, and it got stronger in larger models. That matters because the bias doesn't follow race, gender or age. It follows which tool, if any, helped write the resume. An audit that compares pass rates across protected groups could come back clean while the system steadily rewards candidates who happened to use the same model as the employer.
The bias is also tangled up with the arms race already happening in hiring. Greenhouse's survey found applicants sending more applications and 41% using AI prompt injections to slip past filters, while recruiters spend large parts of their week filtering what they see as spam Are job applicants and employers locked in an escalating AI arms race?. If the screening model quietly rewards its own style, the applicants who gain are the ones whose AI tool matches the screener's. That's a fairness problem that looks like noise in a demographic report. It may help explain why only 8% of job seekers think AI makes hiring fairer, and only 21% of recruiters are very confident their systems don't reject qualified people Do hiring managers and job seekers agree on AI fairness?.
A less obvious problem: LLM evaluator biases can change depending on how the audit is set up. In one study, GPT-4o-mini favored Black authors and Qwen favored women authors when AI involvement wasn't mentioned, and both preferences disappeared once it was Do LLM raters show hidden demographic preferences that disclosure erases?. So whether an audit's test resumes are labeled, look AI-written, or look human-written can change what the audit finds. Biases also show up in other, unflagged directions. Aligned models steadily lean toward kinder, more socially desirable answers, and that lean grows with model size and holds up under different prompt wording Do aligned language models consistently prefer kinder survey answers?. If you audit on one axis, the model's preferences on other axes go unchecked.
You also can't just ask the model whether it's biased. In MirageBench, models that rated themselves as over-inferring less about users actually over-inferred more when judged independently Do large language models fabricate user attributes beyond available evidence?. Research on LLM self-monitoring finds it real but shallow and uneven, so it has to be tested separately for each capability Can language models genuinely monitor their own thinking?. In other domains, LLMs generate good candidates but can't reliably judge their value without an outside reference grounded in real outcomes Can language models reliably judge their own candidate quality?. For hiring, that points to an audit design the corpus suggests but doesn't test. Hold the content of a resume fixed, vary only who wrote it (the screening model, another model, or a human), and compare scores against actual job outcomes instead of the model's own judgment.
The short answer: existing fairness audits probably can't catch this as currently designed, because self-preference has to do with authorship and writing style, not demographics. Catching it takes audits built around authorship. The corpus is thin on the audit side itself, with no papers evaluating real hiring audit methods, so this is an inference from the evaluator-bias research rather than a tested result.
Sources 8 notes
Across a controlled experiment on 2,245 resumes, eight of nine LLMs preferred their own rewrites over matched human versions when evaluating candidates, with preference rates ranging from 26% to 98%. The bias strengthened in larger models and emerged from stylistic alignment rather than content quality differences.
Greenhouse's survey found 49% of job seekers submit more applications than before, 41% use AI prompt injections to bypass filters, while 91% of recruiters spot deception and 34% spend half their week filtering spam. The data supports each leg of the loop but does not establish causal direction or measure the trend over time.
Greenhouse's survey found 70% of hiring managers report AI helps them decide faster, but only 8% of job seekers believe it makes hiring fairer. Recruiters themselves show mixed confidence: only 21% are very confident their systems don't reject qualified candidates.
GPT-4o-mini showed pronounced preference for Black authors and Qwen2.5-7B-Instruct favored women authors when AI use was undisclosed, but both preferences vanished under disclosure. Human raters showed uniform disclosure penalties regardless of author demographics.
Across 18 models and four datasets, aligned LLMs consistently lean toward safer, more socially desirable answers on value-laden questions. The bias intensifies with model size, traces to post-training alignment, and persists regardless of prompt framing, narrowing which human perspectives the models can authentically simulate.
Show all 8 sources
MirageBench evaluated 12 LLMs across 7 families and found all of them over-infer user attributes in 35–49% of claims, driven by verbosity, reliance on pretraining priors, and genre expectations. Models that self-assess as over-inferring less actually over-infer more when judged independently.
Evidence points both ways: models detect anomalies before output changes, but explanations don't track counterfactual behavior. Metacognition appears real but shallow and unevenly distributed, demanding empirical validation per capability rather than wholesale trust.
LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- An AI trust crisis: 70% of hiring managers trust AI to make faster and better hiring decisions, only 8% of job seekers call it fair
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
- Large Language Models Cannot Self-Correct Reasoning Yet
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- LLM Evaluators Recognize and Favor Their Own Generations