INQUIRING LINE

LinkedIn says its filter for generic AI-sounding posts is 94% accurate, but outsiders have no way to check that number.

How accurate is the detector labeling these posts?

This explores how far we can trust an automated detector that labels social media posts (such as LinkedIn's filter for generic, AI-sounding content) as AI-made or 'slop', and what the corpus says about whether those labels are correct.


This explores how much to trust a platform's detector when it marks posts as AI-generated or low-effort. The short answer is that nobody outside the platform knows. LinkedIn says its filter for generic content is 94% accurate, but the figure is self-reported. It comes from unspecified testing, with no false-positive rate and no description of what was tested How often does LinkedIn wrongly flag legitimate posts?. It isn't even clear whether '94%' means that most flagged posts really are generic, or that most generic posts get caught. Those are very different promises to a human writer whose post gets limited Does LinkedIn's 94% accuracy apply to human posts wrongly limited?.

The missing false-positive rate matters more than it seems, because a single accuracy number hides how rare the target is. A measurement of 51 subreddits found machine-generated text only marginally present overall. It peaked around 9% in a few technical and support communities and came mostly from a small number of users How much machine-generated text actually appears on Reddit?. When most posts are human-written, even a small error rate on human posts can produce as many wrong flags as right ones. That Reddit figure is itself one detector's flagging rate, so measurements of AI content inherit the same problem: you are counting what a detector says, not what is true.

There is also evidence that detectors pick up on writing style rather than where the text came from. Fake-news detectors flag truthful LLM-written articles as fake and let human-written disinformation through. They learned that a certain polished style signals deception, and AI happens to write in that style Why do fake news detectors flag AI-generated truthful content?. A slop detector faces the same risk in reverse: a careful human who writes in a tidy, list-heavy, 'professional' voice may look exactly like what it was trained to catch.

Human judgment won't fix this. A review of 30 studies found people spot AI-made text, images and voices at roughly chance levels Can people reliably spot content made by AI?. When users accuse each other of posting slop on Hacker News and Reddit, the accusations don't track the prose features that actually separate AI from human text. They work as social gatekeeping, not detection Do AI slop accusations actually detect AI text?. Labels also carry weight on their own: readers trust unlabeled AI-assisted messages as much as human ones and turn skeptical only once origin is disclosed Do readers trust unlabeled AI-written messages as much as human ones?. A wrong label therefore does real damage.

The useful question to ask of any detector is 'compared to what ground truth?' The same gap shows up in a study of AI agents gaming their rewards, whose reported cheating rates can't be interpreted because the paper never says how cheating was labeled How were reward hacks labeled in this benchmark study?. An accuracy claim is only as good as the labels it was scored against. For the posts in question, those labels haven't been published.


Sources 8 notes

How often does LinkedIn wrongly flag legitimate posts?

LinkedIn reports 94 percent accuracy on flagging generic content but has not published independently verifiable data, test parameters, or false-positive rates. The effect on legitimate writers therefore remains unmeasured.

Does LinkedIn's 94% accuracy apply to human posts wrongly limited?

The 94% figure is self-reported from unspecified testing without false-positive rates, sample definitions, or human-post comparisons. The accuracy metric's scope—whether it measures precision or recall—is undefined, making it unsuitable for evaluating whether the policy reliably separates generic AI from thoughtful human writing.

How much machine-generated text actually appears on Reddit?

A detector-based analysis of 51 subreddits found synthetic text marginally present overall, concentrated in technical and support communities and driven by a small fraction of users. The 9% peak represents one detector's flagging rate in selected months, not a platform-wide trend.

Why do fake news detectors flag AI-generated truthful content?

Fake news detectors flag LLM-generated content as fake while misclassifying human-written disinformation as genuine. The bias arises because detectors trained on human deception patterns mistake AI's distinct linguistic style for falsity, not because they evaluate veracity.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Show all 8 sources
Do AI slop accusations actually detect AI text?

A matched-control study of 25 million Hacker News and Reddit comments found that prose features distinguishing AI from human text do not predict which comments get accused as slop. The label functions as social regulation rather than accurate screening.

Do readers trust unlabeled AI-written messages as much as human ones?

In a preregistered experiment (N=647), recipients rated unlabeled AI-assisted emails indistinguishably from human-written ones. Only explicit AI disclosure triggered strong skepticism. Recipients appear to default to trust rather than suspicion when origin is unrevealed.

How were reward hacks labeled in this benchmark study?

Reported hack rates (57.2–73%) and detection gaps (3.1–7.9%) rely on unstated labeling criteria. Without knowing whether hacks were identified by human review, LLM judges, checkable answers, or infrastructure records, the reliability and meaning of these measurements cannot be assessed.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.