INQUIRING LINE

Slop can be grammatically flawless and still hollow. What concrete signs set it apart from merely clumsy writing?

What observable quality dimensions distinguish slop from other forms of poor writing?

This explores what you can actually point to in a piece of text that marks it as 'slop', as opposed to writing that is just clumsy, ungrammatical or badly organized, and whether those markers have anything to do with whether a machine wrote it.


This explores what makes slop a distinct kind of bad writing rather than a new name for an old problem. The corpus gives a fairly sharp answer: classic poor writing usually fails at the surface, with typos, tangled sentences or broken structure. Slop tends to pass the surface and fail underneath. When researchers coded definitions from 19 experts, slop broke into three axes, each of which can be measured What dimensions make text feel like AI slop?. The first is information utility: how much useful, relevant content there is per sentence. The second is information quality: whether it's factual and how biased it is. The third is style quality, meaning repetition and template-like phrasing. Grammar and correctness don't appear on that list. A text can be clean and still be slop.

That raises a point worth separating out: slop is a judgment about quality, not a test of who wrote it Can we judge text quality without knowing who wrote it?. A padded corporate memo written by a person can be slop, and a tight machine-written summary can avoid being slop. Treating slop this way splits two questions that usually get blurred: what does this text read like, and who produced it? The first dimension can even be counted. Knowledge density measures unique facts or ideas per token. LLM text scores lower than human writing because models elaborate and repeat, adding words without adding content Can we measure reading efficiency as a quality metric?. That gives you a concrete test for padding: a long text that says little.

The second quality that sets slop apart is argumentative flatness. LLMs have mastered grammar and organization but avoid taking evaluative positions. They favor neutral 'manner' nouns ('the approach', 'this process') over the evidential and status nouns human writers use to signal what they think is true or important Why does AI writing sound generic despite being grammatically correct?. The result is prose that is coherent and goes nowhere. Clumsy human writing often has the opposite profile: messy on the surface, but clearly arguing for something. Factual quality can also hide behind fluency. In document-editing tasks, weaker models visibly delete content, while frontier models quietly corrupt it and leave the surface looking intact Does model capability change how documents degrade?.

Here's the twist. The very polish that defines slop also makes readers worse at catching it. Evaluators have rated AI-written documents as both human and better than real human submissions Does polished writing actually signal better quality work?. LLM judges fall for confident formatting and fake citations in the same way Can LLM judges be fooled by fake credentials and formatting?. Newer models are moving further from human word-choice patterns, yet they're getting harder for people to spot, probably because training rewards high quality ratings rather than human-like writing Why do newer AI models diverge further from human writing patterns?. And writers keep about 96% of AI paragraphs unchanged, editing only 23% of the time, so these traits reach readers largely unfiltered Do writers actually edit AI-generated text before publishing?.

The most surprising finding: when people actually call something slop in the wild, they aren't tracking these dimensions. A study of 25 million Hacker News and Reddit comments found that the prose features that separate AI text from human text don't predict which comments get accused Do AI slop accusations actually detect AI text?. So there are two versions of 'slop'. One is a measurable quality profile: low density, little factual reliability, templated style and no stance. The other is a social label used for gatekeeping. The research supports the first. Everyday use mostly reflects the second.


Sources 10 notes

What dimensions make text feel like AI slop?

Coded definitions from 19 experts yield three axes: information utility (density and relevance), information quality (factuality and bias), and style quality (repetition and templatedness). Each axis maps to automatic or human-annotated proxies for assessment.

Can we judge text quality without knowing who wrote it?

Research distinguishes slop—a quality assessment based on coherence and relevance—from AI-text detection, which identifies authorship origin. The framework applies equally to human and machine-written texts, separating what a text reads like from who produced it.

Can we measure reading efficiency as a quality metric?

Knowledge Density (KD) operationalizes reading efficiency by dividing unique atomic knowledge units by text length. LLM-generated text scores lower on KD than human writing because retrieval redundancy and the model's tendency to elaborate inflate token count while holding knowledge content constant.

Why does AI writing sound generic despite being grammatically correct?

AI text uses manner nouns and anaphoric references that are descriptively neutral, while human writers use status and evidential nouns that carry evaluative weight. This produces organizationally coherent but argumentatively inert prose.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Show all 10 sources
Does polished writing actually signal better quality work?

Studies show evaluators perceived AI-generated documents as both human-written and better quality than human submissions. This suggests rhetorical polish misleads judgment and should not serve as a quality signal in evaluation.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Why do newer AI models diverge further from human writing patterns?

ChatGPT-4.5 and o4-mini show greater lexical diversity differences from human text than earlier models, yet human judges cannot reliably distinguish them. Training objectives like RLHF appear to optimize for quality ratings rather than human-like writing patterns.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Do AI slop accusations actually detect AI text?

A matched-control study of 25 million Hacker News and Reddit comments found that prose features distinguishing AI from human text do not predict which comments get accused as slop. The label functions as social regulation rather than accurate screening.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.