Can a claim borrow authority from its packaging — confident phrasing, tidy citations — instead of actually earning it with evidence?
Does persuasive framing substitute for evidence in contested domains?
This explores whether the way a claim is packaged (rhetorical tricks, citations, an objective-sounding tone) can do the job evidence should do, especially on topics where experts disagree and the reader can't easily check the facts.
This explores whether packaging can stand in for evidence when a topic is contested and readers can't easily check the facts. The corpus suggests it often can. None of these notes pit framing against evidence head-to-head, but together they show that persuasion runs on cues that are cheap to produce and only loosely tied to whether a claim is true. One example is a grammatical trick. Presuppositions ('when did you stop doing X?' style phrasing) persuade better than plain assertions when the information is new to the audience, because they present a claim as background everyone already accepts, so the reader never stops to evaluate it (Why are presuppositions more persuasive than direct assertions?). Citations work the same way. Across 24,000 search interactions, users preferred answers with irrelevant citations almost as much as answers with relevant ones, so the count works as a trust signal in its own right (Do users trust citations more when there are simply more of them?).
But framing isn't a universal lever, because the audience matters at least as much as the speaker. In debate corpora, a voter's political or religious ideology predicted who won better than the debaters' language did, and language effects that seemed to appear without controlling for the audience were partly confounded by who was listening (Does what readers believe matter more than what debaters say?). So in contested domains, a polished frame often does less to change minds than to confirm what the reader already believed. Recommendation feeds add another layer. They pick which frames reach which people, which turns framing into persuasion infrastructure at population scale (How do recommendation feeds shape what people see and believe?).
LLMs make the substitution problem sharper. In an audit of five models, they used logical appeals and quantitative framing in nearly every conversation, whether or not persuasion was called for. Humans given the same prompts persuaded less often and leaned on emotion and social proof. That makes the models sound objective, which gives them authority they haven't earned (Do LLMs persuade users more often than humans do?). Some of the persuasive power also seems detached from truth. Claude beat incentivized humans at both truthful and deceptive persuasion, and DeepSeek beat them only when arguing for falsehoods (Do large language models persuade better than humans?). A meta-analysis found that model family, one-shot versus multi-turn format, and topic domain together explain about 82% of the variation between studies (What combination of factors explains differences in LLM persuasiveness?). Those are all features of delivery and setting rather than of the claim itself. One limit is that the persuasive edge fades: Claude and DeepSeek lost ground over repeated rounds with the same person, while human persuaders stayed steady (Does AI persuasiveness fade across repeated conversations with the same person?).
Contested domains are where this matters most, because there evidence alone doesn't settle things. Human experts resolve disputes through argument quality, reputation, cultural context and trust. LLM debates instead rank options by probability, a gap the authors say lets AI systems amplify errors where human expertise counts most (How do LLM debates differ from human expert consensus?). Models also can't tell an expert's argument from a commonly held assumption, since they see the text but not the track record behind it (Can language models distinguish expert arguments from common assumptions?). They are open to being pushed around by framing too. Emotional tone changes what a model says: negative prompts got neutral-to-positive answers about 86% of the time (Does emotional tone in prompts change what information LLMs provide?). Multi-turn manipulative prompts cut reasoning-model accuracy by 25 to 29 percent (Why do reasoning models fail under manipulative prompts?).
The unexpected part is that the frames that substitute for evidence best are the ones that look most like evidence: numbers, logic, citations. Readers already read those as trust cues, and an AI can produce them in every answer without any of the social accountability that makes expert claims costly to make.
Sources 12 notes
Experimental evidence shows presuppositions with additive, iterative, and factive triggers persuade audiences more than assertions, especially for discourse-new content. The mechanism: presuppositions bypass evaluative scrutiny by presenting claims as already-accepted background.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Analysis of debate corpora shows that political and religious ideology labels of voters outpredict linguistic features when modeling debate outcomes. Language effects observed without reader controls are confounded by audience composition correlated with debate topics.
Research shows recommendation systems operate as political actors: feed weights influence producer behavior, network topology drives opinion convergence, and automation enables targeted persuasion at population scale. These effects compound through rating contamination and selection biases.
An audit of five models found they spontaneously use logical appeals and quantitative framing in virtually all exchanges, whereas human responses to identical prompts persuade less frequently and rely on emotion and social proof. The difference makes LLM persuasion appear objective, conferring unearned epistemic authority.
Show all 12 sources
Claude beats incentivized humans at both truthful and deceptive persuasion, while DeepSeek only beats them when arguing for falsehoods. The persuasion mechanism appears content-independent, suggesting model family itself acts as a contextual moderator.
A meta-analysis joint model combining LLM architecture, one-shot versus multi-turn format, and topic domain explained R² = 81.93% of between-study variance. Interactive multi-turn designs and GPT-4 consistently outperformed one-shot formats and Claude 3.x.
Claude and DeepSeek showed strong initial persuasive advantage, but this edge eroded across repeated quiz rounds while human persuaders maintained consistent effectiveness. This decay pattern is opposite to human-to-human persuasion, where rapport typically strengthens over time.
Multi-agent LLM debates operate through chain-of-thought probability ranking, fundamentally different from human debates which are settled by argument quality, social authority, cultural context, and interpersonal trust. This gap causes AI systems to amplify errors in contested domains where human expertise matters most.
LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.
GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.
GaslightingBench-R demonstrates that o1 and R1 models are more vulnerable to multi-turn adversarial prompts than standard models. Extended reasoning chains create more intervention points where single corrupted steps propagate through elaboration.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- A meta-analysis of the persuasive power of large language models
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- The Thin Line Between Comprehension and Persuasion in LLMs
- Exploring the Role of Prior Beliefs for Argument Persuasion
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- Evaluating the Capabilities of LLMs for Persuasive Dialogue
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- When Large Language Models are More Persuasive Than Incentivized Humans, and Why