When ChatGPT pulled people away from Stack Overflow, did it take mostly low-value posts, or good ones too?
What proportion of Stack Overflow's lost posts were genuinely high-quality?
This explores whether the Stack Overflow posts that stopped appearing after ChatGPT launched were mostly low-value material (duplicates, beginner questions) or included genuinely good content, and whether anyone has put a number on that split.
This explores whether the Stack Overflow posts that disappeared after ChatGPT arrived were mostly low-value material or included real high-quality content, and whether that share has been measured. The short answer is that the corpus has no proportion. What it has is a clue: the posts that kept appearing on Stack Overflow didn't get better or worse votes after ChatGPT's release Did ChatGPT displace only low-quality Stack Overflow posts?. If ChatGPT had only absorbed the easy, repetitive, or sloppy questions, the remaining posts should have scored noticeably higher on average. They didn't, which suggests the lost posts looked like a cross-section of the site, good posts included.
That inference has a weak spot. It treats upvotes as a measure of quality, and no expert ever checked whether the displaced posts were actually useful Did ChatGPT displace only low-quality Stack Overflow posts?. Votes reflect visibility, timing, and community habits as much as technical merit. So the finding rules out the reassuring story that 'only the junk left.' It doesn't tell you whether 10% or 60% of what was lost was valuable.
The same gap shows up across the rest of the corpus on AI and online content: quality and authorship claims tend to rest on proxies nobody has independently checked. LinkedIn says it catches generic AI posts with 94% accuracy, but it hasn't published false-positive rates or test details, so nobody knows how many good human posts get suppressed along with the slop How often does LinkedIn wrongly flag legitimate posts? Does LinkedIn's 94% accuracy apply to human posts wrongly limited?. On Reddit, detector-based studies find that the amount of AI text depends heavily on post format and community. Top-level posts are far more likely to be AI-written than replies, and technical and support forums (the communities most like Stack Overflow) show the highest rates Why does Reddit's AI share seem so low compared to others? How much machine-generated text actually appears on Reddit?. That suggests one useful way to sharpen the question: within Stack Overflow, which kinds of posts disappeared, and which kinds survived?
Here's the part you might not have expected: the question is hard to answer for a structural reason, not just a data one. Quality on a Q&A site gets built up after posting, through answers, edits, and votes. A question that was never asked never gets that chance. Some of the lost posts might have become valuable reference threads, and no metric can recover the value of a conversation that never happened. Meanwhile, the open web is filling with AI-assisted text whose meanings are getting more uniform even though measured accuracy holds steady How much of the internet is AI-generated now?. The concern may be less that good posts were lost and more that the range of questions people bring into public view is narrowing.
Sources 6 notes
Vote scores on Stack Overflow showed no significant change after ChatGPT's release, suggesting the displaced content included high-quality posts, not merely duplicates or poor-quality material. However, this conclusion relies on votes as a proxy for quality without expert validation.
LinkedIn reports 94 percent accuracy on flagging generic content but has not published independently verifiable data, test parameters, or false-positive rates. The effect on legitimate writers therefore remains unmeasured.
The 94% figure is self-reported from unspecified testing without false-positive rates, sample definitions, or human-post comparisons. The accuracy metric's scope—whether it measures precision or recall—is undefined, making it unsuitable for evaluating whether the policy reliably separates generic AI from thoughtful human writing.
Reddit's 4.4% aggregate AI share results from replies comprising 72% of scanned items at 98.1% human-authored, while top-level posts were 11.6% AI. Format heavily influences AI rates, with top-level posts 5.25 times more likely to be AI-generated even after controlling for length.
A detector-based analysis of 51 subreddits found synthetic text marginally present overall, concentrated in technical and support communities and driven by a small fraction of users. The 9% peak represents one detector's flagging rate in selected months, not a platform-wide trend.
Show all 6 sources
Internet Archive analysis (2022-2025) shows 35% of newly published websites are AI-generated or AI-assisted. This correlates with declined semantic diversity and increased positive sentiment, but factual accuracy and stylistic diversity remain unchanged.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- AI Content Is Everywhere on Social Media, Especially LinkedIn
- AI Now Writes as Many Online Articles as Humans
- Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- LinkedIn AI Content Study: 81% of Long-Form Posts Are Likely AI
- The Impact of AI-Generated Text on the Internet
- Keeping conversations real on LinkedIn
- LinkedIn's war on AI slop is not just a policy update—it is an admission that the platform lost control of its feed