INQUIRING LINE

Can you tell which paragraphs in a document a chatbot wrote and which a person did, or only software can?

Can readers tell which parts of a document were AI-generated versus human-written?

This explores whether a person reading a document can spot which passages an AI wrote and which a human wrote, and whether anything else can do it instead.


This explores whether a person reading a document can spot which passages an AI wrote and which a human wrote, and whether anything else can do it instead. The corpus's answer is mostly no by eye, yes by machine. No note tests paragraph-level marking inside a single mixed document, so that part is inferred from adjacent findings.

On readers, the evidence is bleak. AI text differs measurably from human text on six vocabulary-diversity dimensions, and newer models diverge further. Yet human judges, including trained linguists, can't reliably tell the difference (Can humans detect AI text if machines can measure it?, Can human judges detect measurable differences in AI text?). Reading passively, which is how most of us read most documents, is worse. In a displaced Turing test, people and AI judges reading transcripts scored below chance. Only an interrogator who could ask follow-up questions kept a marginal edge (Can humans detect AI by passively reading its text?). A document can't be questioned, so that edge is unavailable.

Readers do pick up something, though, without being able to name it. AI-written posts read as aloof because they don't make the internal appeal to the reader's attention that human writing makes (Does AI writing lack the internal appeal to attention that humans use?). AI prose is also grammatically fluent but argumentatively inert, because it avoids the evaluative stance-taking that human writers do (Why does AI writing sound generic despite being grammatically correct?). The effect shows up in how readers judge the author. In a study with 11,091 readers, AI assistance shifted their impression of the writer on all 29 measured dimensions, toward more extreme, confident and agreeable (Does AI writing assistance change how readers perceive the writer?). So the AI's fingerprints reach the reader's impression even when they don't register as AI.

Machines do much better. Simple, interpretable linguistic features caught LLM-written arguments with 99% accuracy (Can simple linguistic features detect AI-written arguments?). A detector looking only at narrative structure, such as character agency and chronological order, separated AI fiction from human fiction with 93.2% accuracy, and it held up with style cues stripped out. These structural choices can't be fixed by polishing the wording, only by rewriting (Can AI stories be detected without analyzing writing style?). One theory of why: artificial text lacks embodied authorship and a real conversational counterpart, so it is structurally missing things human text has, and AI hotel reviews are flagged at 80%+ accuracy because they are false about personal experience (Does AI-generated text lose core properties of human writing?).

This matters for mixed documents because the human touch rarely erases these signals. Writers edited AI-generated paragraphs only 23% of the time, and the edits kept about 96% of the original (Do writers actually edit AI-generated text before publishing?). Writers also chose the AI version of their own paragraph 63% of the time, even though it distorted their stance (Do writers actually prefer AI-edited versions of their own text?). So AI passages tend to reach readers nearly intact, and readers can't flag them. The corpus points to provenance as the practical fix: one system binds every number, quote and asset to its source, making traceability, not fluency, the trust test (Can source traceability make AI writing trustworthy?). Since perception fails, a document would have to carry its own record of where each part came from.


Sources 12 notes

Can humans detect AI text if machines can measure it?

LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.

Can human judges detect measurable differences in AI text?

Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.

Can humans detect AI by passively reading its text?

The displaced Turing test shows that both human and AI judges reading transcripts performed below chance accuracy, while interactive interrogators retained marginal detection ability. The adaptive advantage of real-time questioning collapses entirely in passive consumption.

Does AI writing lack the internal appeal to attention that humans use?

Human writing contains an appeal to the reader's attention as a fundamental property of communication itself. AI-generated posts inherit platform visibility but do not perform this internal appeal, producing the reported aloofness readers perceive — a structural absence, not a stylistic defect.

Why does AI writing sound generic despite being grammatically correct?

AI text uses manner nouns and anaphoric references that are descriptively neutral, while human writers use status and evidential nouns that carry evaluative weight. This produces organizationally coherent but argumentatively inert prose.

Show all 12 sources
Does AI writing assistance change how readers perceive the writer?

A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.

Can simple linguistic features detect AI-written arguments?

General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Does AI-generated text lose core properties of human writing?

Research shows artificial text disrupts dialogic symmetry, context continuity, embodied authorship, and political situatedness. These are not surface flaws but structural absences—AI hotel reviews show 80%+ detection accuracy due to inherent falsity about personal experience distinct from human deception.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Do writers actually prefer AI-edited versions of their own text?

In a study of 4,503 cases, 63% of writers chose AI-generated text over their own original paragraphs, with 52% claiming the AI version better reflected their views. This preference persisted across three AI models despite evidence that AI versions systematically distort the original stance.

Can source traceability make AI writing trustworthy?

Data2Story's Inspector binds every number, quote, and asset to its origin, making provenance rather than fluency the adoption gate. Across 18 samples, human raters favored this approach, showing that verifiable derivation—not surface polish—enables professional newsrooms to adopt agent output.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.