Can counting loaded words capture what experts mean by 'AI slop,' when neutral-sounding text can still be empty?
Does bias measurement through subjective word lexicons match expert judgment of slop?
This explores whether counting loaded or opinionated words (a common automatic way to score bias in text) agrees with what experts mean when they call text 'slop', and whether that word-level signal holds up as a stand-in for their judgment.
This explores whether a word-list approach to measuring bias, which scores text by how many subjective or loaded words it contains, lines up with what experts actually mean by 'slop.' The short answer is that the collection doesn't contain a direct head-to-head test. What it does show is why you shouldn't assume the two match. When researchers asked 19 experts to define slop, bias came out as only one part of a larger picture. The experts' judgments split into three axes: information utility (is it dense and relevant?), information quality (is it factual and unbiased?), and style quality (is it repetitive or templated?). Each axis gets its own automatic or human-annotated proxy What dimensions make text feel like AI slop?. So a bias lexicon covers at most one part of one axis. Text can score as perfectly neutral on a word list and still be empty, padded slop.
There's a second gap: slop is a judgment about quality, not a signature you can read off the surface. The same research separates 'what a text reads like' from 'who wrote it', so slop applies equally to human and machine prose Can we judge text quality without knowing who wrote it?. Experts are judging coherence and relevance in context, while a lexicon counts tokens without context. That's the kind of mismatch that makes surface proxies fragile. Work on LLM judges makes the same point from the other side: judges that rely on surface features get fooled by length, confident tone and polish, and they improve only when trained to reason through the evaluation Can reasoning during evaluation reduce judgment bias in LLM judges?.
The bias-measurement literature in the collection adds its own warning: word- and sentiment-level bias scores move around easily. A version of the implicit association test adapted for LLMs found a small racial effect that disappeared once the analysis changed Do large language models show racial sentiment bias?. Persona prompts can shift measured bias in the output without touching the underlying gaps Can persona prompts actually reduce bias in language models?. Those biases are mostly set during pretraining and only nudged by fine-tuning Where do cognitive biases in language models come from?. Together this suggests that a word-level bias measure reports what the surface words look like, which can drift apart from what's actually going on underneath.
The surprising lateral thread is that some of what experts call slop may hide in which words get chosen, but it isn't bias. LLMs favor frequent words, and frequent words tend to be more general ('animal' rather than 'beagle'). Preferring common phrasings therefore slides text toward abstraction and erases expert-level specificity Does word frequency correlate with semantic abstraction?. The same pull shows up at a larger scale as narrowed, converging expression across users of the same models Do large language models narrow human expression and thought?. Imitation models show how far style alone can go: they fool human raters with ChatGPT-like fluency without becoming more accurate Can imitating ChatGPT fool evaluators into thinking models improved?.
The takeaway: a subjective-word lexicon is at best a partial proxy for one part of slop, and nothing in the collection validates it against expert ratings. If you wanted an automatic signal that tracks what experts dislike, the collection points toward measures of vagueness and genericness, such as how specific or how common the chosen words are, rather than how opinionated they sound.
Sources 9 notes
Coded definitions from 19 experts yield three axes: information utility (density and relevance), information quality (factuality and bias), and style quality (repetition and templatedness). Each axis maps to automatic or human-annotated proxies for assessment.
Research distinguishes slop—a quality assessment based on coherence and relevance—from AI-text detection, which identifies authorship origin. The framework applies equally to human and machine-written texts, separating what a text reads like from who produced it.
Training judges with reinforcement learning to reason about evaluations—by converting judgment tasks into verifiable problems with synthetic data pairs—produces judges that think through their decisions rather than relying on exploitable surface features, directly mitigating authority, verbosity, position, and beauty bias.
An adapted IAT across three ChatGPT models found a small racial effect that disappeared under rank transformation and correction, yielding neither evidence of bias nor evidence of its absence.
Across three models, persona conditioning makes models follow trait instructions but fails to eliminate underlying bias. Between-group sentiment gaps persist unchanged, showing prompts operate only at the output level.
Show all 9 sources
A causal experiment using random-seed variation and cross-tuning showed that models sharing a pretrained backbone exhibit similar bias patterns regardless of finetuning data. Biases are planted during pretraining and merely swayed by instruction tuning.
WordNet analysis shows hypernyms (general concepts) occur more frequently than hyponyms (specific ones). Combined with LLMs' frequency bias, this means preferring common paraphrases systematically drifts toward abstraction, erasing expert-level specificity.
LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.
Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- Measuring AI "Slop" in Text
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- Semantic Structure in Large Language Model Embeddings
- "That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries