Can admissions readers tell AI essays from human ones, and what happens to real applicants they get wrong?
What error rates do admissions officers have when identifying AI writing?
This explores how accurately admissions officers can tell AI-written application essays from human-written ones: how often they catch AI, how often they wrongly flag a real applicant, and what those mistakes cost.
This explores how accurately admissions officers spot AI-written essays and how often they get it wrong. The short answer is that the collection doesn't give a clean error rate. No study here reports a false-positive or false-negative percentage for admissions readers. What it does have is a picture of why that number matters, and of why human detection is shakier than it feels to the people doing it.
The closest evidence is an experiment in which admissions officers could *often* tell AI essays from human ones, and gave lower ratings to essays they believed were AI-written Do admissions officers penalize essays they suspect are AI-written?. The rating hinges on that belief, not on whether the essay really was AI-written. That makes errors costly in both directions: a human applicant whose essay reads as 'AI-ish' gets marked down anyway. The real-world stakes show up in a study of 7,500 applications to a public policy master's program. Most 2025 applicants submitted likely-AI essays despite a ban, and they were admitted at lower rates than comparable applicants, even though AI improved essay quality Does AI essay use hurt admissions chances despite quality gains?. The authors suggest reader suspicion as the reason, but that link is proposed, not measured.
Research outside admissions should make anyone cautious about trusting a reader's instinct. AI text differs measurably from human text on several measures of word variety, yet human judges, trained linguists included, can't reliably detect it. Newer models are drifting *further* from human style while getting *harder* to spot Can humans detect AI text if machines can measure it?. A study of 25 million Hacker News and Reddit comments found that the features that actually distinguish AI prose don't predict which comments get accused of being 'slop.' The accusation works more as social gatekeeping than as detection Do AI slop accusations actually detect AI text?. If an admissions reader's suspicion behaves the same way, they may be reacting to a style (flat, organized prose that never takes a stance Why does AI writing sound generic despite being grammatically correct?) that some human writers share and that edited AI text can avoid.
The bigger issue is what admissions essays were for in the first place. A simulation of Freelancer.com hiring found that removing writing as a signal made hiring 19% less meritocratic: strong workers got hired less often and weak ones more often, because writing had served as proof of effort Does cheap writing weaken hiring based on worker ability?. Admissions officers who penalize suspected AI may be trying to keep that signal alive, but with an instrument (their gut) whose accuracy nobody here has measured. Even openly disclosed AI help draws a small but consistent penalty from both human and AI raters Does disclosing AI assistance make readers trust articles less?, so the bias against 'AI-touched' writing seems to run deeper than catching cheaters.
If you came looking for a number: the gap in the research is the finding. The collection shows that officers act on suspicion and that applicants pay for it. It does not show how often that suspicion lands on the wrong person.
Sources 7 notes
In an experiment, admissions officers could often discriminate AI from human essays and rated essays they believed to be AI-generated lower than those believed human-written. The authors frame this as a plausible explanation for the observed admissions penalty, though the link remains proposed rather than directly measured.
Among 7,500 applications to a public policy master's program, majority of 2025 applicants submitted AI-generated essays despite explicit prohibition. These applicants were admitted at lower rates than similar applicants without detected AI use, despite AI improving essay quality.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
A matched-control study of 25 million Hacker News and Reddit comments found that prose features distinguishing AI from human text do not predict which comments get accused as slop. The label functions as social regulation rather than accurate screening.
AI text uses manner nouns and anaphoric references that are descriptively neutral, while human writers use status and evidential nouns that carry evaluative weight. This produces organizationally coherent but argumentatively inert prose.
Show all 7 sources
A simulation of Freelancer.com hiring without written signals shows top-quintile workers get hired 19% less often, while bottom-quintile workers get hired 14% more often. Employers lose the costly-effort signal that once distinguished able workers.
Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- AI-written admissions essays are widespread but penalized
- Measuring AI "Slop" in Text
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Do LLMs produce texts with "human-like" lexical diversity?
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries