When you mix your own writing with AI's, how often do AI-text detectors wrongly call it machine-made, and has anyone measured that?
What false-positive rates do AI detectors show on mixed human-AI drafts?
This explores how often automated AI-text detectors wrongly flag writing as machine-made when a person drafted or edited it alongside an AI, and what the collection says about that kind of mislabeling.
This explores how often AI-text detectors wrongly accuse writers when a draft mixes human and AI work. The short answer is that the collection has no study measuring detector false-positive rates on blended drafts. Treat any specific percentage you've seen elsewhere as something to check against its source. What the collection does have is the surrounding evidence, and it explains why that number is so hard to pin down and why it matters more than headline accuracy suggests.
Start with the human baseline. A review of 30 studies found that people tell AI-made content from human content at roughly coin-flip accuracy, across text, images, and voice. Their accuracy hasn't improved as AI output has become more realistic (Can people reliably spot content made by AI?). That's why institutions turn to automated detectors in the first place. It also means there's no reliable human check on what a detector decides. A mixed draft makes this worse, because the 'correct' label is unclear even before any tool is involved. If a person outlines, a model drafts, and the person rewrites, there is no clean ground truth to measure false positives against.
The most useful idea for thinking about detector error rates comes from a note that isn't about detection at all. It argues that a system with 95% accuracy, used at scale, still wrongly convicts thousands of people, and that impressive accuracy figures can hide basic statistical mistakes (Can AI models be truly free from human bias?). The same arithmetic applies to detectors. When most submissions are honest, even a small false-positive rate means many of the people flagged are innocent. Mixed drafts fall exactly where a detector's confidence means least. A related note explains why: AI output is inherently changeable, shifting with sampling, prompt wording, and context (Why does AI output change with every prompt and context?). That makes it hard to find a stable 'AI signature' that survives human editing.
Two neighboring threads show what's at stake. In hiring, applicants use AI and prompt tricks to get past filters while recruiters spend large parts of their week screening out what they see as spam. That's an escalating loop in which filtering errors fall on real people (Are job applicants and employers locked in an escalating AI arms race?). And AI gatekeeping can be uneven across groups. One study found guardrails refusing requests at different rates depending on the user's apparent age, gender, or ethnicity (Do AI guardrails refuse differently based on who is asking?). That study is about refusals, not detectors, but it raises a fair question: are a detector's false positives spread evenly across writers, or do they cluster on certain styles of writing?
The collection also points to a different design. Instead of a detector that hands down a verdict, a system can highlight what a human reviewer should look at and leave the judgment with that person. In the decision-making studies behind this idea, that approach reduced anchoring bias (Can AI guidance reduce anchoring bias better than AI decisions?). Another note finds that we still lack ways to measure whether an AI system's errors stay visible and can be challenged (How can we measure whether AI errors stay visible and recoverable?). For a writer falsely flagged on a draft they co-wrote with AI, that is the real gap: the error rate is unknown, and there is often no way to see or contest the error.
Sources 7 notes
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Research shows that 'theory-free' AI models mask bigotry behind high accuracy metrics while committing fundamental statistical errors. A 95% accurate criminal justice system would wrongly convict thousands, demonstrating that model sophistication does not validate causal inference.
AI outputs exhibit essential mutability—they vary with sampling, prompt wording, and audience interpretation. This is not a defect but a defining feature of tokens as media, making them fundamentally different from fixed commodities and resistant to traditional quality assurance.
Greenhouse's survey found 49% of job seekers submit more applications than before, 41% use AI prompt injections to bypass filters, while 91% of recruiters spot deception and 34% spend half their week filtering spam. The data supports each leg of the loop but does not establish causal direction or measure the trend over time.
GPT-3.5 refuses requests at different rates for younger, female, and Asian-American personas, and sycophantically declines to engage with political positions users would disagree with. Sports fandom and other non-political signals also shift refusal sensitivity.
Show all 7 sources
Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.
Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
- Emergent Introspective Awareness in Large Language Models
- An AI trust crisis: 70% of hiring managers trust AI to make faster and better hiring decisions, only 8% of job seekers call it fair
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- The Return of Pseudosciences in Artificial Intelligence: Have Machine Learning and Deep Learning Forgotten Lessons from Statistics and History?