When an algorithm does the screening, can AI-written text win just by looking polished and authoritative?
Does model-generated text style create systematic advantages in algorithmic screening?
This explores whether text written by AI (or polished with it) gets an unfair edge when another algorithm, such as an LLM judge, grader or filter, does the reviewing, and whether surface style is what tips the scales.
This explores whether AI-styled text gets an unfair edge when an algorithm does the screening, and whether surface style is what wins. The corpus has no direct studies of screening résumés, admissions essays or grant applications. It does have strong evidence on the mechanism such an advantage would run through: automated evaluators reward how text looks, sometimes more than what it says.
The clearest evidence is about LLM judges. Two of their biases are 'semantics-agnostic': an authority bias (trusting text that cites references, even fake ones) and a beauty bias (preferring rich formatting like headers, bullets and bold). Adding fake citations or extra formatting reliably raises a judge's score, with no access to the model and no clever optimization needed (Can LLM judges be fooled by fake credentials and formatting?). Model-generated text tends to come with exactly these features already built in: tidy structure, confident framing, reference-like phrasing. So the advantage doesn't have to be designed. It can show up by default whenever one model's habits match another model's preferences.
Why would model text carry those habits so consistently? One reason: generation follows a smooth path toward the most typical continuation rather than exploring competing positions (Does LLM generation explore competing claims while producing text?). The result is polished, middle-of-the-distribution prose. That is also the kind of text a model-based judge, trained on similar data, finds most 'normal' and credible. A related warning applies: model output reflects the model's learned patterns, not evidence about the world, so treating it as a real signal of quality is a category error (Should we treat LLM outputs as real empirical data?).
The less obvious part is that style is the easiest layer to fake and the easiest to detect, but it isn't the most telling one. Telling authors apart by stylistic patterns is nearly solved, even for small models (Can language models truly understand literary style?). Yet AI fiction stays detectable with style removed entirely, from deeper choices like how characters act and how time is ordered (Can AI stories be detected without analyzing writing style?). For screening, that cuts both ways. A screener that scores surface polish can be gamed. A screener built on deeper structure is sturdier, and it may also flag AI-assisted writing that surface 'humanizing' edits don't hide. There is a related risk: frontier models damage documents through subtle corruption that leaves the surface intact (Does model capability change how documents degrade?). A polished document is not necessarily a sound one, and a style-sensitive screener will struggle to tell the two apart.
In short, the corpus supports the mechanism: automated evaluators can be swayed by formatting and authority cues that AI text produces naturally. It does not yet measure how large that advantage is in real hiring, admissions or peer-review pipelines. That gap is worth noticing in itself.
Sources 6 notes
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Token prediction trains models to continue toward the training distribution, not to explore logically related counterpositions. This smoothness in process produces smooth claims that multiply without generating new perspectives.
Foundation Priors framework shows that LLM-generated text reflects the model's learned patterns and user's prompt choices, not ground truth. Such outputs should only influence inference through explicitly parameterized trust weights, not be treated as equivalent to real evidence.
GPT-2 achieves 95% accuracy identifying authorship through style patterns alone, but lacks the evaluative framework to explain why those stylistic choices carry meaning. Detection without interpretation remains cataloguing, not criticism.
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
Show all 6 sources
DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Argument Collapse: LLMs Flatten Long-Form Public Debate
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- The Assistant Erased You: Measuring Loss of Authorship Signals in AI-Mediated Communication
- StoryScope: Investigating idiosyncrasies in AI fiction
- Humans or LLMs as the Judge? A Study on Judgement Biases