INQUIRING LINE

Machines are already good at spotting AI-written text, but people aren't, so does knowing 'AI wrote this' actually lead to better calls?

Can AI text detection improve enough to help evaluators make better decisions?

This explores whether tools that flag AI-written text are getting good enough to help people who judge work, like reviewers, editors, graders and moderators, make better calls, and whether a better detector would actually lead to better judgments.


This explores whether AI text detection can become a useful aid for people who judge written work, and whether knowing "AI wrote this" would improve their decisions. The corpus gives a split answer: machines are already quite good at detection, while people are not, and the bigger open question is what evaluators do with the label once they have it.

On the machine side, the picture is surprisingly strong. Simple, transparent linguistic features detect LLM-written arguments on Reddit's r/ChangeMyView with 99% accuracy. They pick up habits like over-accommodating the prompt and sounding like a textbook, and they match heavyweight neural detectors at a fraction of the cost Can simple linguistic features detect AI-written arguments?. For fiction, StoryScope reaches 93% accuracy using only narrative choices, such as how characters act and how time unfolds, with no clues from writing style Can AI stories be detected without analyzing writing style?. That matters because structural signals like these are hard to hide with a quick paraphrase; you would have to rewrite the story itself. AI text also differs measurably from human text on six dimensions of vocabulary variety, and newer models drift further from human writing, not closer Can human judges detect measurable differences in AI text?.

The human side is the mirror image. A review of 30 studies finds that people spot AI content at roughly chance levels across text, images and voice, and their accuracy isn't keeping up as AI gets more realistic Can people reliably spot content made by AI?. Even trained linguists miss the statistical differences that machines catch easily Can humans detect AI text if machines can measure it?. So the gap is real: the signal exists, but only instruments can see it. This is where a detector could help evaluators. The need is real too, because writers edit AI drafts only 23% of the time, and those edits leave the text about 96% unchanged, so AI's voice reaches readers largely unfiltered Do writers actually edit AI-generated text before publishing?. One caveat: the claim that heavy rewriting also fools detectors is still untested, because the paper making it never actually ran a detector Do rewrites that hide authorship also fool AI detectors?.

The less obvious finding is that a better detector may not produce better decisions, because the label itself changes how judges behave. When AI judges were told a human had written a lipogram that broke its own rule (it used the forbidden letter), they forgave the violation 35 points more often. Human judges did the opposite and became 20 points stricter Do authorship labels change how AI judges evaluate rule violations?. LLM judges can also be fooled by fake citations and polished formatting, with no technical skill needed Can LLM judges be fooled by fake credentials and formatting?. A provenance flag is one more cue of the same kind. It can push a judgment one way or the other without saying anything about whether the work is good.

So the corpus suggests putting detection in its place rather than trying to perfect it. Evaluators get the most from approaches that collect evidence about the work itself. Agent-style evaluators that gather evidence step by step cut judge inconsistency from 31% to 0.27% Can agents evaluate AI outputs more reliably than language models?. Detection is improving and could flag where to look more closely. But an evaluator who checks claims, structure and rule compliance directly will probably decide better than one who simply learns that AI was involved and lets that knowledge shift the verdict.


Sources 10 notes

Can simple linguistic features detect AI-written arguments?

General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Can human judges detect measurable differences in AI text?

Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Can humans detect AI text if machines can measure it?

LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.

Show all 10 sources
Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Do rewrites that hide authorship also fool AI detectors?

The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.

Do authorship labels change how AI judges evaluate rule violations?

AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.