Measuring AI "Slop" in Text
AI “slop” is an increasingly popular term used to describe low-quality AIgenerated text, but there is currently no agreed upon definition of this term nor a means to measure its occurrence. In this work, we develop a taxonomy of “slop” through interviews with experts in NLP, writing, and philosophy, and propose a set of interpretable dimensions for its assessment in text. Through span-level annotation, we find that binary “slop” judgments are (somewhat) subjective, but such determinations nonetheless correlate with latent dimensions such as coherence and relevance. Our framework can be used to examine AI-generated text in both detection and binary preference tasks, potentially offering new insights into the linguistic and stylistic factors that contribute to quality judgments. We highlight that fully automated and scalable methods remain an open challenge.
Introduction. “Slop” has emerged as a term describing generic, low-quality content that appears to have been generated by AI.1 Recent news articles offer salient examples of such AI “slop”, ranging from nonfactual claims (“... add nontoxic glue to make cheese stick to a pizza”, “geologists advise eating at least one rock a day”; Hoffman, 2024; Wallace-Wells, 2024) to useless information (“fodder for websites whose only purpose appears to be optimising for [search engines]”; Mahdawi, 2025). Conversations on social media highlight indicators of “slop” in LLM responses, including overuse of certain terms, low information density, and structural quirks such as lists-as-responses.2 Despite the sudden ubiquity of the term, there is no clear definition of, nor method, for measuring “slop” in text.
This gap matters: large-scale surveys, such as Microsoft’s Occupational Implications of Generative AI (Tomlinson et al., 2025) and Anthropic’s Economic Index (Handa et al., 2025) reveal AI is primarily used in writing and information gathering tasks. Defining and measuring “slop” may help characterize and ultimately improve LLM writing. Some individuals deeply familiar with AI generated content can reliably detect AI writing on the basis of structural and lexical quirks, even without training (Chakrabarty et al., 2024; Russell et al., 2025). Yet text can be perceived as “slop” even when not generated by AI, and not all AI-generated text reads as “slop”.
Our primary aim in this work is to characterize qualities of texts that contribute to them being categorized as “slop.” Such factors may explain instances where humans mistakenly characterize human-written text as AI-generated, and “slop” might provide an explainable metric that accounts for binary preferences between texts collected from human annotators. We apply principles from measurement theory to conceptualize and operationalize a definition of “slop” (Bandalos, 2018). We aim to provide language for articulating style and components of bothersome LLM-generated text, while also providing a framework for measuring such aspects.
Our main contributions are as follows: We first introduce a working definition and taxonomy of “slop” and map each dimension to automatic metrics where possible (§3). To validate this framework, we collect span-level annotations from expert writers over 150 news articles and 100 question-answering passages to provide a fine-grained analysis of slop indicators (§4). Although
Related work. AI-Text Detection. There is now a small body of work on discriminating between human- and AI-written texts, e.g., DetectGPT (Mitchell et al., 2023) and Binoculars (Hans et al., 2024) provide scores for the likelihood that they were AI generated, and report high discriminant performance (0.95 AUROC). Russell et al. (2025) provide an interpretable guide listing key indicators of AIwritten text. While related, recognizing “slop” differs from AI-text detection in general, and can be applied to any text source (whether AI-written or not). In this work our taxonomy and annotations diverge from those used for AI-text detection in general.
Text Diversity. Prior work has sought to characterize aspects of texts related to how repetitive and templated they are. Salkar et al. (2022) investigated repeated n-grams in LLM outputs in the context of summarization. Shaib et al. (2024b) found that modern LLMs are prone to repeatedly generate favoured syntactic templates, i.e., sequences of Part-of-Speech (PoS) tags. Padmakumar & He (2024) and Tevet & Berant (2020) examined lexical and semantic diversity in generated texts, introducing metrics to quantify variation across outputs and emphasizing its importance to generation quality. These existing efforts have informed the way in which we are thinking about what characterizes writing style and AI “slop” and provides automatic measurements for key aspects of “slop.”
Text Quality Measurements. Text quality has typically been measured using simple surface-level metrics like BLEU (Papineni et al., 2002) and ROUGE (Lin, 2004), which can be effective when reference outputs are available (e.g., machine translation evaluation). More recent work has recognized that text quality is not monolithic but rather comprises multiple, sometimes competing dimensions that must be measured independently, and accordingly focused on multidimensional frameworks assessing properties of texts. Chakrabarty et al. (2025b) provide an editing taxonomy to correct (Chakrabarty et al., 2025a) recurring AI-writing flaws such as cliches and unnecessary exposition. Similarly, Bharadwaj et al. (2025) show that reward models over-weight 5 superficial writing cues including length, structure, jargon, sycophancy, and vagueness. Both works confirm that multiple factors contribute to text quality. Our work is complementary to measuring quality in general: We target stylistic patterns unique to LLM writing that are not covered by other taxonomies.
Method. The Oxford Dictionary defines “slop” as: “[...] material produced using a large language model (LLM), which is often viewed as being low-quality or inaccurate. This type of low-quality, AIgenerated material is becoming increasingly visible to people [...], who often view it as unwanted or inferior.”
“Slop” as a construct does not immediately permit measurement: It is difficult to quantify “lowquality” or “unwanted” text. We propose a composite measure over observable characteristics of text, where we elicit salient characteristics from a set of individuals with a range of relevant expertise. Human writing can also read as “slop”, but we adopt the above definition and focus on (seemingly) LLM-generated texts.3 We first solicited detailed definitions of “slop” from 19 individuals with a range of expertise across relevant disciplines including writing, journalism, linguistics, NLP, and philosophy (App. Table 4b). This group included PhD students, professors, and industry professionals from the listed disciplines. All but one respondent had 3 or more years experience in their field at the time of their response (App. Figure 4a). We asked individuals to describe their familiarity with the term “slop” in the context of AI-generated content, as well as a description of typical use (if any) of LLMs in their work. 11 experts (58%) had encountered the term “slop” as relates to AI-generated content. Most reported using LLMs more than 2 times a week (n = 14). The rest mostly used them sporadically (n = 4) with 1 expert never using them. We asked experts to provide a definition and list key characteristics of text that make it “slop.” We provide the full survey sent to experts in Appendix A.
Using qualitative content analysis and deductive coding techniques, we map expert definitions of “slop” to measurable concepts (Hsieh & Shannon, 2005). We begin by identifying key terms in survey responses and building a code list until saturation (i.e., until no new codes are created). We then map each response on to one of the following codes: Factuality, Information Density, Bias, Relevance, Repetition, Templatedness, Verbosity, Word Complexity, Tone, Coherence, Fluency, Diversity, Engagement, Vagueness, and Utility.
Assigned codes were separately reviewed by all authors, as were disagreements and redundant codes. The codebook was iteratively updated throughout this process. Redundant codes (e.g., Vagueness and Information Density) were collapsed. We further categorize codes with overarching categories or themes: Information Utility, Information Quality, and Style Quality. Table 1 describes the full code hierarchy within each theme, and the count of responses containing each code tag.
Here, we describe each theme and code after annotator adjudication (See Appendix I for a description with examples).
Information Utility assesses how effectively a text conveys meaningful and contextually appropriate information. It comprises two key indicators: (i) Density, or the amount of substantive content relative to the length of the text, measured through information-theoretic token entropy (Meister et al., 2021) and propositional idea density (Brown et al., 2008), and (ii) Relevance, the alignment of content with task or prompt, measured through expert human annotations due to complexities in automated assessments (Clarke & Dietz, 2024).
Information Quality describes the accuracy and subjectivity of the presented information. Factuality assesses inaccuracies, hallucinations, or fallacious claims within the text, which require human annotations due to the complexity of automated factual evaluations in the absence of reference texts (Ramprasad et al., 2024). Bias (Subjectivity) assesses the presence or absence of a necessary subjective or rhetorical perspective, measured by the proportion of subjective words through an established lexicon (Wiebe et al., 2004).
Style Quality addresses properties related to expression and readability. Repetition, identified by lexical repetition metrics (Shaib et al., 2024a) and Templatedness, measured via syntactic structures (Shaib et al., 2024b) are key features of text Structure.
Discussion. LLMs are often deployed as cheap alternatives for human preference judgments in alignment and evaluation (Bharadwaj et al., 2025), however our findings highlight important limitations. Unlike reasoning tasks where rewards are verifiable, for subjective tasks there is a significant risk of miscalibration. Prior work has documented these issues: Chakrabarty et al. (2025a) and Gooding et al. (2025), for example, show that LLMs struggle to select high-quality writing actions as judged by human experts, often treating suboptimal and optimal interventions as equally acceptable. This leads to low quality text that is often referred to as “slop".
A recent study from OpenAI (Chatterji et al., 2025b) shows that almost 50% of ChatGPT usage focuses on writing (28.1%) and information seeking (21.3 %) tasks. To ensure better alignment in such areas, we present the first systematic attempt at qualitatively characterizing “slop” in LLM-generated text. Our findings suggest that Information Quality, Information Utility, and Style Quality are important axes of text evaluation. Further, granular codes within each axis can vary in strength based on the domain, or the purpose of the text. We show that our taxonomy provides a useful framework for assessing writing across domains, beyond accuracy- or reference-based metrics. While overall “slop” judgments are somewhat subjective, our analysis shows that an increase in issues along these axes increases the likelihood of a text being judged as “slop”. We also show that current evaluation practices are not sufficient for automatically measuring “slop”. Existing automatic text metrics fail to capture whether generated text is genuinely useful or well-written relative to “slop”. Neither LLMs-as-judges nor linear models trained over these features are able to fully approximate human assessments of “slop,” however we hope the taxonomy can guide future improvements of LLMbased reward models. While our analysis focuses on text written in English, we hope the framework introduced here can support future analyses of other languages.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can readers reliably distinguish AI-written text from human writing?- What prose features actually separate AI text from human writing?
- Can lexical repetition and syntactic structure predict whether text feels templated?
- Do measurable differences exist between AI text and human writing?
- Do texts judged as slop actually contain measurable stylistic patterns unique to LLMs?
- What observable quality dimensions distinguish slop from other forms of poor writing?
- What prose features distinguish automatically generated text from human writing?
- Does AI-generated writing feel polished while remaining harder to understand?
- Why does production time matter to the meaning of generated text?
- Why does AI writing seem more competent and informative than human writing?
- What signals of individual identity become unreliable in AI-assisted text?
- What structural difference exists between AI posts and human conversational writing?
- Does AI writing erase markers of non-native English speaker identity?
- When do readers defer to AI text without genuine processing?
- Do writers recognize when AI text misrepresents their actual stance?
- Can a classifier distinguish machine-written text from poor human writing?
- Can AI text detectors reliably identify AI-generated websites?
- Why can't algorithms distinguish between human and AI generated content quality?
- Why do human judges fail to detect systematic linguistic differences that classifiers easily identify?
- Can readers detect when text was written or heavily influenced by AI?