Can software reliably tell AI writing from human writing, or does it only work in narrow cases that edits can break?
Can AI detectors reliably distinguish human from machine-generated text?
This explores whether software, not people, can reliably tell AI-written text from human writing, and what the corpus says about where that works and where it breaks.
This explores whether automated detectors, not human readers, can reliably tell machine-written text from human writing. The corpus gives a split answer. In narrow, well-defined settings, machines can detect AI text very accurately. The corpus has little evidence that a general-purpose detector stays reliable once the text has been edited, disguised, or taken out of the setting it was trained on.
Start with the gap between people and machines. Humans are close to useless at this. A review of 30 studies found that human accuracy sits around chance for text, images, and voice, and it hasn't improved as AI output has become more realistic Can people reliably spot content made by AI?. Statistical measures, though, show that AI text really is different. ChatGPT's writing differs from human writing on six separate measures of vocabulary use, including how varied, how evenly spread, and how repetitive its word choices are. Trained linguists still can't see those differences Can human judges detect measurable differences in AI text?. Newer models drift further from human patterns while becoming harder for people to spot Can humans detect AI text if machines can measure it?. Reading a transcript passively is especially weak: both human and AI judges scored below chance, while people who could question the other party live kept a slight edge Can humans detect AI by passively reading its text?. So the signal exists, but only measurement can pick it up.
When detectors are built for a specific kind of text, they do well. Simple, transparent linguistic features caught LLM-written counter-arguments on Reddit's r/ChangeMyView with 99% accuracy, as well as heavyweight neural detectors did. They worked because LLMs leave habits behind, such as echoing the prompt and producing textbook-style argument markers Can simple linguistic features detect AI-written arguments?. The fiction result is more surprising. A system called StoryScope separated AI stories from human ones with 93% accuracy without looking at writing style at all. It used only story-level choices, such as how characters act and how events are ordered in time. Because those choices sit deeper than word choice, a 'humanizer' tool that polishes the prose doesn't remove them. You would have to rewrite the story Can AI stories be detected without analyzing writing style?. This suggests the most durable detection signals may lie in how AI organizes ideas, not in its vocabulary.
There are two weak spots. The first is that detectors can learn the wrong lesson. Fake-news detectors flag truthful AI-written articles as fake and let human-written disinformation through. They treat 'sounds like an LLM' as if it meant 'deceptive,' so they are detecting style, not truth Why do fake news detectors flag AI-generated truthful content?. The second is that rewriting is the obvious way to fool a detector, and the corpus hasn't tested it. One paper claims heavily rewritten messages slip past AI detectors, but it never runs a detector. It only shows that the rewrites' style converges Do rewrites that hide authorship also fool AI detectors?. The current gap may favor detectors for now: writers edit AI-drafted paragraphs only 23% of the time, and their edits leave about 96% of the text unchanged Do writers actually edit AI-generated text before publishing?. Most AI text reaches readers with its fingerprints intact.
The overall picture: AI text is statistically distinct from human writing, people can't see the difference, and specialized machines usually can. 'Reliable' still depends on staying in the genre the detector was trained on, on nobody deliberately disguising the text, and on not confusing 'AI-written' with 'false.' The corpus doesn't test commercial detectors against determined paraphrasing, so it can't yet answer the hardest version of this question.
Sources 9 notes
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
The displaced Turing test shows that both human and AI judges reading transcripts performed below chance accuracy, while interactive interrogators retained marginal detection ability. The adaptive advantage of real-time questioning collapses entirely in passive consumption.
General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.
Show all 9 sources
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
Fake news detectors flag LLM-generated content as fake while misclassifying human-written disinformation as genuine. The bias arises because detectors trained on human deception patterns mistake AI's distinct linguistic style for falsity, not because they evaluate veracity.
The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
- Do LLMs produce texts with "human-like" lexical diversity?
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts
- Measuring AI "Slop" in Text
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship