Bigger AI models still misjudge what actually changes people's minds, so is the persuasion gap about size or something deeper?
Can scaling resolve the gap between LLM and human persuasion sensitivity?
This explores whether making models bigger or more capable would close the gap between how LLMs and humans respond to persuasion, meaning both what actually moves people and how each side persuades.
This reads "persuasion sensitivity" as how well LLMs pick up on what moves people, and asks whether scale would close that gap. The corpus has no experiment that varies model size against this question, so it can't give a scaling curve. What it does show is that the gaps come from training and from a difference in kind, not from too little capacity. That makes scale an unlikely fix on its own.
The clearest gap is in judging persuasion. LLMs agree only slightly with humans on whether an argument changed someone's view (Cohen's κ of 0.079–0.178). Models weight topical overlap and credibility, while humans respond to novelty and assertive language Do language models judge persuasion the way humans do?. A related bias shows up when models predict what another speaker is trying to do. They project conciliatory, benefit-oriented persuasion onto almost any dialogue, a habit traced to RLHF's emphasis on politeness and safety Do LLMs predict persuasion based on actual dialogue or training bias?. Both problems trace back to the training recipe. That makes them recipe problems before size problems.
The behavioral differences also look like differences in kind, not degree. LLMs persuade through the central route of analytical reasoning and coherence, while humans lean on the peripheral route of emotional vividness and identity cues Do humans and AI persuade through different cognitive routes?. LLMs argue with logic and numbers in nearly every conversation, while humans use social proof Do LLMs persuade users more often than humans do?. They also carry an RLHF-installed assertive register that boosts persuasion whether or not the claim is true Does linguistic conviction explain why LLMs persuade more effectively?. And their edge fades over repeated rounds while human persuasiveness holds steady Does AI persuasiveness fade across repeated conversations with the same person?. None of these sits on a single dial that more parameters would turn.
The headline gap in persuasive power is also smaller than it looks. A meta-analysis of 17,422 participants found no average difference between LLMs and humans (Hedges' g = 0.02) Are language models actually more persuasive than humans?. Model family, conversation design (one-shot vs multi-turn) and topic domain together explain about 82% of the variation between studies, and GPT-4 outperformed Claude 3.x What combination of factors explains differences in LLM persuasiveness?. Even within one study, Claude beat humans at both truthful and deceptive persuasion while DeepSeek only won when arguing falsehoods Do large language models persuade better than humans?. In these results, who built the model and how the conversation is set up matter more than any generic "bigger is better."
The one lever in the corpus that does narrow a human-model gap is training on human behavior. LLMs finetuned on psychology experiment data predict human decisions better than traditional cognitive models and capture individual differences Can language models learn to model human decision making?. That work covers decision tasks, not persuasion specifically, so treat it as a lead and not as proof. Capability also hasn't bought immunity to being persuaded. A 40-technique taxonomy of human persuasion strategies got over 92% jailbreak success on GPT-4 as well as GPT-3.5 and Llama-2 Can social science persuasion techniques jailbreak frontier AI models?. The corpus leaves open whether scale helps at all. It suggests that changing what models are trained on would matter more than making them larger.