Can readers tell truth from fabrication without evidence signals?
When readers see fluent text with no provenance information, do they distinguish accurate claims from AI-generated hallucinations? This tests whether presentation authority alone misleads judgment.
The paper reports that readers given no provenance signal could not separate truth from fabrication. In the control condition, high-fluency hallucinations averaged M = 6.28 and ground truth M = 5.78 (p = .43). The excerpt does not state the rating scale behind those means. When an idealized Provenance Density interface was shown, the authors report a discernment gap of +4.15 points (d = 1.82, p < .001) between truth and fabrication. They call the control result a "Fluency Trap": readers trust fluent text as if fluency still cost something to produce. Participant comments fit that reading. One said the claims "sound believable and well written," and another judged accuracy by "how the text was worded and flowed."
The authors' mechanism runs through costly signaling. Polish once worked as a handicap, since faking it was expensive, so it separated competent writers from others. Generative AI collapses that separating equilibrium into a pooling one, where expert testimony and fabrication share the same form. Provenance Density is their response: a score built from atomic claims, each checked by retrieval and weighted by source reputation and relevance, displayed as the density of verified claims. Binary "Made with AI" labels answer "Who wrote this?" rather than "What supports this?" The excerpt says binary labels "acted as a blunt warning," and a participant called the banner "a warning vs being informative." The authors read the label as a risk cue, an accuracy-independent discount they call the Transparency Penalty. The excerpt gives no statistics for the label condition.
This extends Does polished AI output trick audiences into trusting it? from charts and decks to prose, and tests the reader side. The presentation authority that note describes is what the no-signal participants appear to have leaned on. It also gives the fluency illusion in How do AI tools trick users into overestimating their own skills? a measured counterpart. That note concerns a producer's inflated self-assessment, while this paper shows fluency standing in for truth in a reader's judgment. The paper also sits beside Do users worldwide trust confident AI outputs even when wrong?. That note locates the problem in confidence cues. The paper tests a verified-claim cue and does not test confidence markers, so it cannot say whether evidence density would correct that overreliance.
The result is an upper bound by design. The Oracle protocol showed participants idealized values, with grounded summaries paired with high-density indicators and fabricated ones with null indicators, so the user study does not test the live pipeline. The technical audit (N = 200 on a composite of TruthfulQA and FreshQA) found that "retrieval density alone is insufficient," and the authors list retrieval reinforcing popular misconceptions as a live risk. The Latin-square design did not fully cross interface, veracity and topic; the Provenance Density–hallucinated cell was measured on one topic, Matcha. The score is also conservative by design: established facts averaged 0.79 and emerging dynamic topics 0.64. The evidence supports that a correct evidence-density display can change how people discriminate. It does not show that a deployed pipeline produces correct densities, or that the effect holds across topics. Treat provenance density as a tested direction for transparency, with high-density false positives as the open risk.
Inquiring lines that read this note 28
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does AI-generated content create social proof without authentic interaction? Can readers reliably distinguish AI-written text from human writing?- Why does an hour of coffee count as harder to fake than written prose?
- Can style and perspective be reliably separated by automated detection systems?
- How much does polished presentation substitute for actual expertise in reader judgment?
- Does polished text presentation hide process-level authenticity from readers?
- Can audience attention alone explain why disclosure triggers stronger skepticism?
- Can transparency about how and when AI was used rebuild reader trust?
- Why does suspicion of AI origin trigger skepticism but not complete dismissal?
- What false positive rate do citation verification tools produce on archival works?
- How much undetected fraud exists beyond current retraction statistics?
- Can statistical detection of synthetic text identify actual fraudulent manuscripts?
- How do courts distinguish between AI hallucinations and ordinary typographical errors?
- Does performing the source verification work create meaningful engagement with ideas?
- Can checking someone else's proof count as genuine mathematical understanding?
- Does publishing proofs without showing the verification process undermine mathematics?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
same mechanism, style standing in for expertise; this paper tests it with readers of prose.
-
How do AI tools trick users into overestimating their own skills?
When people use language models to help with work, what system-level properties create false confidence in their own competence? Understanding this matters for recognizing hidden skill gaps.
the fluency illusion there concerns producers' self-assessment; here fluency substitutes for truth in readers.
-
Do users worldwide trust confident AI outputs even when wrong?
Explores whether the tendency to over-rely on confident language model outputs transcends language and culture. Understanding this pattern is critical for designing safer human-AI interaction across diverse linguistic contexts.
that note finds confidence cues dominate; this paper tests an evidence cue and not confidence markers.
-
Can agreement across samples reveal when models are wrong?
This explores whether sampling a model multiple times and checking consistency can catch false answers. It matters because consistency checks are often used as safety measures, but may have blind spots.
the gate the audit credits with most of the signal, and its stated blind spot.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Verbal lie detection using Large Language Models
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
- Beyond Hallucinations: The Illusion of Understanding in Large Language Models
- Seeing to Think? How Source Transparency Design Shapes Interactive Information Seeking and Evaluation in Conversational AI
- "That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
Original note title
participants given no signal showed no detectable truth discernment — idealized Provenance Density restored it in an 81-person study