Do AI-text detectors truly spot machine authorship, or do they just flag clumsy writing, even when a human wrote it?
Can a classifier distinguish machine-written text from poor human writing?
This explores whether AI-text detectors are actually spotting machine authorship, or just flagging bad writing, and so would wrongly catch a clumsy human writer too.
This explores whether AI-text detectors are actually spotting machine authorship, or just flagging bad writing, and so would wrongly catch a clumsy human writer too. The most useful idea in the corpus is that these are two different questions. 'Slop' is a judgment about quality: is the text coherent, relevant and worth reading? Detection asks something else: who or what produced it? Slop can come from a human or a machine (Can we judge text quality without knowing who wrote it?). When experts break slop down, they find three things: how much useful information it carries, whether that information is accurate and fair, and stylistic problems like repetition and templated phrasing (What dimensions make text feel like AI slop?). A careless human essay can fail all three without any AI involved.
The detectors that work best don't seem to be looking for 'badness' at all. They pick up on habits that are oddly polished. One classifier reached 99% accuracy on AI-written Reddit counter-arguments. It did this by spotting the way LLMs mirror the prompt and use textbook argument markers, habits typical of a model that writes too neatly, not of a weak writer (Can simple linguistic features detect AI-written arguments?). AI prose is grammatically confident but avoids taking an evaluative stance. It is well organized but doesn't argue for anything (Why does AI writing sound generic despite being grammatically correct?). Its vocabulary also follows measurable statistical patterns across six dimensions of word variety and spread (Can humans detect AI text if machines can measure it?). In fiction, AI shows up in how the story is built, such as how much agency characters have and how time is ordered, even after all surface style is removed (Can AI stories be detected without analyzing writing style?). Poor human writing usually goes wrong in a different way. It tends to be messy, uneven and opinionated, while AI text tends to be smooth, even-handed and inert. So in principle, a good classifier should be able to tell them apart.
People can't do this reliably. Across 30 studies, human detection of AI content sits around chance (Can people reliably spot content made by AI?). Judges also let labels do the work: the same passage gets rated higher when it's marked 'human', and AI evaluators show this bias 2.5 times more strongly than people do (Do authorship labels bias how we judge literary quality?). This suggests that when a person calls a piece 'AI-sounding', they may often mean 'bad' or 'generic', which is exactly the confusion the question is about.
The corpus has a real gap here. None of these studies tests detectors against deliberately poor human writing, such as rushed students, non-native writers or SEO filler. So the claim that classifiers can tell 'bad human' from 'machine' is an inference from what the detectors attend to, not a measured result. The boundary is also blurring from both sides. Writers edit AI drafts only 23% of the time and change very little when they do, so a lot of 'human' text is really lightly touched AI (Do writers actually edit AI-generated text before publishing?). Heavy rewriting may wipe out authorship signals entirely, though that claim hasn't been tested against actual detectors yet (Do rewrites that hide authorship also fool AI detectors?). The surprising takeaway is that machine text is easiest to spot when it's at its most fluent. A detector built well should flag the suspiciously competent paragraph rather than the clumsy one.
Sources 10 notes
Research distinguishes slop—a quality assessment based on coherence and relevance—from AI-text detection, which identifies authorship origin. The framework applies equally to human and machine-written texts, separating what a text reads like from who produced it.
Coded definitions from 19 experts yield three axes: information utility (density and relevance), information quality (factuality and bias), and style quality (repetition and templatedness). Each axis maps to automatic or human-annotated proxies for assessment.
General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.
AI text uses manner nouns and anaphoric references that are descriptively neutral, while human writers use status and evidential nouns that carry evaluative weight. This produces organizationally coherent but argumentatively inert prose.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
Show all 10 sources
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Human judges rated identical passages 13.7 percentage points higher when labeled human-authored; AI models showed a 2.5-fold stronger bias at 34.3 points. The effect persists across AI architectures, suggesting evaluators respond to provenance cues rather than text quality alone.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Measuring AI "Slop" in Text
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- Do LLMs produce texts with "human-like" lexical diversity?