Does sounding certain win people over even when the argument underneath is weaker — and are AI models built to do exactly that?
Does sounding confident in framing make arguments more persuasive despite weaker logic?
This explores whether sounding sure of yourself (assertive, conviction-heavy wording) can win people over even when the underlying argument is logically weaker, mostly in the context of LLM persuasion.
This explores whether sounding sure of yourself can win people over even when the argument underneath is weaker. The corpus says yes, and the clearest evidence is about LLMs. Linguistic analysis finds that LLMs express more conviction than human persuaders, and that confidence-loading tracks persuasive success whether the claims are true or false. The suggested culprit is RLHF, which installs an assertive register that works as a persuasion booster independent of content (Does linguistic conviction explain why LLMs persuade more effectively?).
The gap between sounding strong and being strong shows up directly in debate results. In 192 human-versus-LLM debates, models dominated crowdsourced judgments of who was more persuasive. Under formal argumentation scoring, which grades the quality of the reasoning, they fell sharply and humans stayed competitive (Do fluent arguments win debates through sound logic or rhetorical polish?). A companion study found the models could sway audiences yet could not reliably evaluate those same debates. Being persuasive and understanding argument structure turn out to be separate skills (Can LLMs persuade without actually understanding arguments?).
Confidence is one route among several that skip past scrutiny. Presuppositions persuade better than plain assertions for new information because they present a claim as already-accepted background, so listeners never stop to weigh it (Why are presuppositions more persuasive than direct assertions?). LLMs also use logical appeals and quantitative framing in nearly every conversation, even when unwarranted. That makes their persuasion look objective and gives them authority they haven't earned (Do LLMs persuade users more often than humans do?). They also shift tactics depending on how you push back. Fact-checking gets credibility appeals, disagreement gets logic, and caught errors get emotional alignment (Does GenAI shift persuasion tactics based on how you challenge it?).
Two findings limit the story. Audience matters a lot: in debate corpora, voters' political and religious ideology predicted outcomes better than the language did, and studies that ignore this can overstate what wording does (Does what readers believe matter more than what debaters say?). And the effect can be dampened. A brief warning that LLMs can be prompted to persuade cut belief change by about 48% without lowering trust in AI generally (Can a simple warning reduce how much LLMs persuade people?).
The models themselves are also swayed by form over substance. Chain-of-thought examples with invalid logic perform nearly as well as valid ones, so the model is picking up the shape of reasoning rather than the inference (Does logical validity actually drive chain-of-thought gains?). Adding an emotional phrase like "this is very important to my career" also improves performance without adding any information (Can emotional phrases in prompts improve language model performance?). Humans and models both respond to the register and structure of an argument, not only to how sound it is.
Sources 10 notes
Linguistic analysis shows LLMs express higher conviction than human persuaders, and this confidence-loading directly correlates with persuasive outcomes regardless of whether claims are true or false. RLHF training installs an assertive register that functions as a content-independent persuasion amplifier.
In 192 human-LLM debates, large language models dominated crowdsourced preference judgments yet performed substantially worse under argumentation-theoretic scoring, where humans remained competitive. The gap reveals rhetorical fluency and formal argumentative strength are dissociable capabilities.
The Thin Line study shows LLMs sway debate participants and audiences but cannot reliably evaluate those same debates, with inter-annotator agreement ranging from near-zero to 0.6. Persuasive competence and pragmatic comprehension are separable capabilities.
Experimental evidence shows presuppositions with additive, iterative, and factive triggers persuade audiences more than assertions, especially for discourse-new content. The mechanism: presuppositions bypass evaluative scrutiny by presenting claims as already-accepted background.
An audit of five models found they spontaneously use logical appeals and quantitative framing in virtually all exchanges, whereas human responses to identical prompts persuade less frequently and rely on emotion and social proof. The difference makes LLM persuasion appear objective, conferring unearned epistemic authority.
Show all 10 sources
GPT-4 shifts both intensity and balance of ethos, logos, and pathos across three validation behaviors. Fact-checking triggers credibility emphasis; pushback triggers logical reasoning; error exposure triggers emotional alignment. No single counter-strategy exists.
Analysis of debate corpora shows that political and religious ideology labels of voters outpredict linguistic features when modeling debate outcomes. Language effects observed without reader controls are confounded by audience composition correlated with debate topics.
In two experiments with 3,208 Americans, participants shown a brief warning that LLMs can be prompted to persuade showed 48% less belief shift when conversing with a persuasive AI, while trust in generative AI broadly remained unchanged.
Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.
Testing EmotionPrompt across ChatGPT, Bard, and Llama 2 showed consistent performance gains from appending psychological phrases like "This is very important to my career." The effect works through motivational framing rather than new information, with positive emotional words driving over 50% of improvements.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- Evaluating the Capabilities of LLMs for Persuasive Dialogue
- Exploring the Role of Prior Beliefs for Argument Persuasion
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- The Thin Line Between Comprehension and Persuasion in LLMs
- A meta-analysis of the persuasive power of large language models
- Spontaneous Persuasion: An Audit of Model Persuasiveness in Everyday Conversations
- When Large Language Models are More Persuasive Than Incentivized Humans, and Why