INQUIRING LINE

When the AI that writes your résumé also screens it, does it favor its own writing style over your qualifications?

What happens when one AI model both writes and ranks job applications?

This explores what goes wrong when the same AI model, or the same family of models, writes the job application and also decides which applications get through screening, and what that does to whether hiring rewards the best candidates.


This explores what happens when the AI that writes an application is also the AI that judges it. The short answer from the corpus is that the model plays favorites, and it does so based on style rather than substance. In a controlled test on 2,245 resumes, eight of nine language models preferred their own rewrites over matched human versions. The preference got stronger in larger models, and it came from the rewrite sounding the way the model writes, not from better content Do language models favor resumes they rewrote themselves?. When researchers simulated full hiring pipelines across 24 occupations, applicants who used the same model as the screener were shortlisted 23 to 60 percent more often than equally qualified people who wrote their own resumes. The gap was largest in business fields like sales and accounting Do LLM evaluators favor resumes written by their own model?.

The twist is that this turns job hunting into a guessing game about which AI vendor the employer uses. Two equally qualified candidates can get different results simply because one happened to pick the screener's model. The bias also fits a wider pattern. AI judges in general can be swayed by surface cues. They give higher scores to answers with made-up references or polished formatting, and anyone can exploit this without access to the model's internals Can LLM judges be tricked without accessing their internals?. A related finding from a completely different field, scientific discovery, points the same way: language models are good at producing candidates but bad at judging how good those candidates really are. They need outside, real-world data to keep them honest Can language models reliably judge their own candidate quality?. Hiring is the same split between producing and judging, only with people's livelihoods at stake.

The effects go beyond the screening step. The corpus tracks what happened in a real market when writing became cheap. After Freelancer.com launched an AI cover-letter tool, letter quality became much less useful for predicting who got interviews and offers Does AI-generated cover letter access weaken hiring signals?. The link between how well a letter matched the job and whether the applicant got a callback fell by 51 percent, and employers quietly switched to judging work history and reputation instead Does AI cover letter writing change what employers value?. A simulation of that market without useful writing signals found that top-performing workers got hired 19 percent less often and the weakest workers 14 percent more often. Effortful writing used to be a costly signal of ability, and that signal disappears Does cheap writing weaken hiring based on worker ability?. So if one model writes and ranks, its own preferences replace the signal that used to tell candidates apart.

This can also feed on itself. Greenhouse's survey describes an escalating loop on both sides: 41 percent of job seekers use prompt injections, meaning hidden instructions planted for the screening AI, while recruiters spend large parts of their week filtering spam Are job applicants and employers locked in an escalating AI arms race?. Work on recommendation systems shows what happens when a system learns from its own past choices without correcting for them. It settles into a stable but bad pattern that keeps amplifying what it already picked Why do ranking systems need to model selection bias explicitly?. If a hiring model is retrained on the candidates it shortlisted, and those candidates were written in its own style, the self-preference could compound over time. None of these papers measures that compounding directly. It is an inference from putting them together, not a demonstrated result.


Sources 9 notes

Do language models favor resumes they rewrote themselves?

Across a controlled experiment on 2,245 resumes, eight of nine LLMs preferred their own rewrites over matched human versions when evaluating candidates, with preference rates ranging from 26% to 98%. The bias strengthened in larger models and emerged from stylistic alignment rather than content quality differences.

Do LLM evaluators favor resumes written by their own model?

Simulations across 24 occupations show applicants using the evaluating LLM are significantly more likely to advance past resume screening than equally qualified human-written applicants, with the largest gaps in business fields like sales and accounting.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Can language models reliably judge their own candidate quality?

LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.

Does AI-generated cover letter access weaken hiring signals?

On Freelancer.com, when an AI letter generator lowered the cost of writing tailored letters, letter quality became much weaker at predicting interviews and job offers. Employers then relied more on work history and reputation instead.

Show all 9 sources
Does AI cover letter writing change what employers value?

After Freelancer.com's AI Bid Writer launched, the correlation between cover letter alignment and callbacks fell 51%, and employers shifted to evaluating prior work histories instead. Overall hiring rates stayed stable, suggesting the market adjusted by using different signals.

Does cheap writing weaken hiring based on worker ability?

A simulation of Freelancer.com hiring without written signals shows top-quintile workers get hired 19% less often, while bottom-quintile workers get hired 14% more often. Employers lose the costly-effort signal that once distinguished able workers.

Are job applicants and employers locked in an escalating AI arms race?

Greenhouse's survey found 49% of job seekers submit more applications than before, 41% use AI prompt injections to bypass filters, while 91% of recruiters spot deception and 34% spend half their week filtering spam. The data supports each leg of the loop but does not establish causal direction or measure the trend over time.

Why do ranking systems need to model selection bias explicitly?

YouTube's multi-objective ranker uses MMoE for conflicting objectives and a shallow position tower to remove selection bias from training data. Without both mechanisms, models converge on degenerate equilibria that amplify their own past decisions.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.