INQUIRING LINE

If nobody tells an AI who wrote a text, can it spot machine-written text, and how does it compare with people?

Can language models detect AI-generated text in blind evaluation tasks?

This explores whether machines, including language models, can tell AI-written text from human-written text when nobody tells them which is which, and how that compares with human judges.


This explores whether machines can tell AI-written text from human writing when the source is hidden. First, a caveat about the collection: it has strong material on purpose-built detectors and on human judges, but nothing that directly tests a general-purpose LLM asked to make this call blind. What it does show is that the signal is there to be found, that humans mostly miss it, and that LLM judges have weak spots that should make you cautious about using one as the judge.

The most surprising finding is the gap between what can be measured and what people can perceive. AI text differs from human text in measurable ways across six kinds of vocabulary variety. Yet human judges, trained linguists included, can't reliably spot the difference. Newer models drift further from human patterns while getting harder for people to catch Can humans detect AI text if machines can measure it?. Machines don't need to rely on intuition. Simple, transparent linguistic features reached 99% accuracy flagging LLM-written counter-arguments on Reddit's r/ChangeMyView, matching heavyweight neural detectors Can simple linguistic features detect AI-written arguments?. The giveaways weren't odd word choices. The AI arguments leaned too closely on the wording of the prompt and read like textbook examples of a good argument.

The signal also goes deeper than style. In fiction, AI stories could be separated from human ones with 93% accuracy using only storytelling choices, such as how much agency characters have and whether events are told in order. Even with every stylistic cue removed, the method kept 97% of its accuracy Can AI stories be detected without analyzing writing style?. This matters for evasion: you can polish surface style with a quick paraphrase, but changing narrative structure means rewriting the story. One paper claims that heavy rewriting also fools AI detectors, but it never actually tests a detector, so that claim is still open Do rewrites that hide authorship also fool AI detectors?. In practice, most AI text never gets rewritten at all. Writers edited AI paragraphs only 23% of the time, and the edited versions stayed 96% similar to the original Do writers actually edit AI-generated text before publishing?. So detectable patterns usually reach readers intact.

There is a twist if the judge is itself an LLM. LLM judges reliably fall for fake references and rich formatting, and these tricks need no access to the model Can LLM judges be fooled by fake credentials and formatting?. A judge that is swayed by how text looks is a fragile choice for a task where the best signals sit below the surface. The stakes are real: one demonstration produced 288 complete finance papers, with invented theories and fabricated citations, from a set of statistically significant signals Can AI generate hundreds of fake academic papers automatically?.

The takeaway: AI text can be detected, but mostly by measuring features people don't notice, such as vocabulary spread, echoes of the prompt and story structure. Asking a model, or a person, whether a piece 'feels' AI-written works much less well. Whether a general LLM asked blind can do as well as these targeted feature methods is a gap in the collection.


Sources 7 notes

Can humans detect AI text if machines can measure it?

LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.

Can simple linguistic features detect AI-written arguments?

General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Do rewrites that hide authorship also fool AI detectors?

The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Show all 7 sources
Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Can AI generate hundreds of fake academic papers automatically?

A demonstration showed LLMs generating 288 complete finance papers from 96 statistically significant signals, each with invented theoretical justifications and fabricated citations, proving academic HARKing can be automated at scale.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.