What dimensions make text feel like AI slop?
Can we break down the vague notion of AI slop into measurable components? Researchers coded expert definitions to find which specific text properties people associate with low-quality generated writing.
The paper treats AI "slop" as a construct to decompose rather than a verdict. It adopts the Oxford Dictionary's definition of low-quality, LLM-produced material, then collects written definitions from 19 people across writing, journalism, linguistics, NLP and philosophy. The abstract calls these responses interviews; the method section describes survey responses. Deductive coding of those responses yields three axes. Information Utility combines Density, measured with token entropy and propositional idea density, and Relevance, judged by human annotators. Information Quality combines Factuality, which the paper says needs human annotation, and Bias (subjectivity), measured as the proportion of subjective words in a lexicon. Style Quality is read through Repetition (lexical repetition metrics) and Templatedness (syntactic structure).
The authors argue that slop "does not immediately permit measurement," because "low-quality" and "unwanted" are hard to quantify. They therefore propose "a composite measure over observable characteristics of text," elicited from people with relevant expertise, and offer the axes as a framework for "assessing writing across domains, beyond accuracy- or reference-based metrics." The paper also says granular codes "can vary in strength based on the domain, or the purpose of the text." The excerpt does not give the final code count after redundant codes were collapsed, or the tallies in its Table 1.
Against the nearest notes, this is a sharper version of the style-versus-substance argument. Does polished AI output trick audiences into trusting it? warns that presentation can stand in for expert judgment; the taxonomy keeps Style Quality as one axis beside Information Quality and Utility rather than letting it carry the whole judgment. Can imitating ChatGPT fool evaluators into thinking models improved? shows a related split in a different setting, with imitation matching surface style while the factuality gap stays open. The density axis sits near Can we measure reading efficiency as a quality metric?, though it uses token entropy and propositional idea density rather than an atomic-unit count. The paper's finding that neither LLM judges nor linear models "fully approximate" human slop assessments fits How much does rhetorical style shift AI review scores?, which shows LLM reviewers responding to rhetorical presentation.
The excerpt does not report the span-level annotation results behind its headline claim. The abstract says binary slop judgments are "(somewhat) subjective" but correlate with latent dimensions such as coherence and relevance; the excerpt gives no agreement statistic, correlation coefficient or annotation protocol, and the 150-article study appears only in the contributions list. The taxonomy rests on a small survey, all but one of whose respondents had three or more years' experience in their field. The implication is that the axes are a usable design space for critiquing text, while automatic scoring of slop should wait for the evidence the excerpt leaves out.
Inquiring lines that read this note 9
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can readers reliably distinguish AI-written text from human writing?- What prose features actually separate AI text from human writing?
- Can lexical repetition and syntactic structure predict whether text feels templated?
- Do measurable differences exist between AI text and human writing?
- Do texts judged as slop actually contain measurable stylistic patterns unique to LLMs?
- What observable quality dimensions distinguish slop from other forms of poor writing?
- What prose features distinguish automatically generated text from human writing?
- Does AI-generated writing feel polished while remaining harder to understand?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
presentation standing in for substance; the taxonomy keeps style as one axis beside information quality.
-
Can imitating ChatGPT fool evaluators into thinking models improved?
Explores whether fine-tuning weaker models on ChatGPT outputs creates an illusion of capability gains. Investigates why human raters and automated judges fail to detect that imitation improves style but not underlying factuality or reasoning.
a related style-versus-factuality split, measured in a different setting.
-
Can we measure reading efficiency as a quality metric?
How can we quantify whether generated text delivers novel information efficiently or wastes reader attention through redundancy? This matters because standard coherence and fluency scores miss texts that are well-written but informationally dense.
neighboring density measure; this paper uses token entropy and propositional idea density instead.
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
fits the finding that LLM judges do not fully approximate human slop assessments.
-
Can we judge text quality without knowing who wrote it?
Does the concept of 'slop' work as a quality judgment independent of whether a machine or human authored the text? This matters because current AI detection often conflates two separate questions: origin and quality.
sibling: the quality framing this taxonomy is built to measure.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Measuring AI "Slop" in Text
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
- AI Skills Improve Job Prospects: Causal Evidence from a Hiring Experiment
- "That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
Original note title
AI slop decomposes into three axes — information utility, information quality and style quality — each mapped to proxies