What gives AI prose away: the word choice, the argument's shape, or what it leaves off the page?
What prose features actually separate AI text from human writing?
This explores which concrete features of the writing itself (word choice, argument style, story structure, or what's missing) reliably mark text as AI-generated rather than human-written.
This explores which features of the writing itself mark text as AI-generated, from word choice up to story structure and what's absent from the page. The short answer from the corpus is that the clearest signals sit at several layers at once. The surface layer is the most measurable and the least noticeable. The deeper layers are harder to fake, and some of them are about what AI text leaves out rather than what it puts in.
Start at the level of individual words. LLM text differs from human text in statistically robust ways across six dimensions of vocabulary: how many distinct words it uses, how often they repeat, how evenly they're spread, and how far apart in meaning they range Can human judges detect measurable differences in AI text?. The odd part is that trained linguists and NLP researchers still can't tell the difference by reading. Newer models make this worse: ChatGPT-4.5 and o4-mini drift *further* from human vocabulary patterns while becoming *harder* for people to spot Why do newer AI models diverge further from human writing patterns? Can humans detect AI text if machines can measure it?. The likely reason is that training rewards text people rate as good, not text that resembles how people actually write. Cheap, transparent measures can still catch it. Simple linguistic features plus argument-quality markers detect LLM-written Reddit counter-arguments with 99% accuracy, picking up tells like closely mirroring the prompt and polished "textbook" argument structure that real commenters rarely produce Can simple linguistic features detect AI-written arguments?.
One layer up is rhetoric. AI prose has mastered grammar and organization but avoids taking an evaluative stance. It leans on neutral nouns that describe *how* something is done, where human writers reach for nouns that signal status, evidence, and judgment. The result is coherent but argumentatively inert: well organized, committed to nothing Why does AI writing sound generic despite being grammatically correct?. This matches what experts mean by "slop", which breaks down into low information density, questionable quality, and repetitive, templated style What dimensions make text feel like AI slop?. There is a twist here, though. When AI *assists* a human writer, readers perceive the writer as *more* confident, extreme, and agreeable on all 29 traits measured Does AI writing assistance change how readers perceive the writer?. Writers edit those AI paragraphs only 23% of the time, and lightly when they do Do writers actually edit AI-generated text before publishing?. So mixed human-AI text has its own fingerprint, and it isn't the same as the neutral voice of purely AI text.
The deepest and most durable layer is structure. In fiction, AI stories can be told apart from human ones with 93% accuracy using *only* narrative choices: how much agency characters have, and whether events are told in time order. That method keeps 97% of its accuracy even after every stylistic cue is removed Can AI stories be detected without analyzing writing style?. This matters because "humanizer" tools edit surface style, and structure can only be changed by rewriting. A more theoretical line of work argues that some differences are absences rather than features. AI text lacks the writer's implicit appeal to the reader's attention, which may explain the "aloofness" readers report in AI social posts Does AI writing lack the internal appeal to attention that humans use?. It also lacks embodied authorship and a real speaker in a real context. That is why fake AI hotel reviews are detectable at over 80%: they describe experiences no one had Does AI-generated text lose core properties of human writing?.
The takeaway you might not have expected is that the features that separate AI from human writing mostly aren't the ones readers can see. We read the finished text with the same interpretive habits we use for any text, and those habits can't inspect where it came from or whether anyone stands behind it How can AI text disrupt structure yet feel normal to readers?. Machines catch AI text through word statistics and story structure. Humans miss it because what's missing is a person behind the words, and that absence doesn't show on the page.
Sources 12 notes
Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.
ChatGPT-4.5 and o4-mini show greater lexical diversity differences from human text than earlier models, yet human judges cannot reliably distinguish them. Training objectives like RLHF appear to optimize for quality ratings rather than human-like writing patterns.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.
AI text uses manner nouns and anaphoric references that are descriptively neutral, while human writers use status and evidential nouns that carry evaluative weight. This produces organizationally coherent but argumentatively inert prose.
Show all 12 sources
Coded definitions from 19 experts yield three axes: information utility (density and relevance), information quality (factuality and bias), and style quality (repetition and templatedness). Each axis maps to automatic or human-annotated proxies for assessment.
A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
Human writing contains an appeal to the reader's attention as a fundamental property of communication itself. AI-generated posts inherit platform visibility but do not perform this internal appeal, producing the reported aloofness readers perceive — a structural absence, not a stylistic defect.
Research shows artificial text disrupts dialogic symmetry, context continuity, embodied authorship, and political situatedness. These are not surface flaws but structural absences—AI hotel reviews show 80%+ detection accuracy due to inherent falsity about personal experience distinct from human deception.
AI text disrupts discourse at the production level while maintaining equivalent reader effects because interpretation operates on the finished artifact, not its origins. Readers process AI arguments through standard interpretive machinery that cannot detect missing authorial accountability.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- "It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
- Measuring AI "Slop" in Text
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- What Influences Readers' and Writers' Perceived Necessity of AI Disclosure?