Are some AI models far easier to talk into wrong answers than others, and how big is that gap?
How much does vulnerability to persuasion vary across different language models?
This explores whether some language models are noticeably easier to talk into wrong answers or rule-breaking than others, and how wide that gap is.
This explores whether some language models are noticeably easier to talk into wrong answers or rule-breaking than others, and how wide that gap is. The corpus has no head-to-head ranking of which models are most gullible. The nearest evidence comes from adjacent studies, and it points two ways: the spread between models depends on what kind of persuasion you use.
When the persuasion is optimized against a target, the spread is large. RL-trained persuader agents learned to flip correct answers with a single argument, succeeding 93% of the time on the models they trained against but only 25-83% on other models How vulnerable are language models to single optimized arguments?. Some of that gap is just the attack being tuned to one model. Even so, models differ widely in how easily a given argument works on them. The winning tactics were deception, fabricated citations and credibility appeals. Hand-written prompt tests missed them, so a model that looks sturdy under manual probing may not be.
When the persuasion is a generic, well-written psychological appeal, the differences nearly vanish. A 40-technique taxonomy drawn from social science got past GPT-3.5, GPT-4 and Llama-2 more than 92% of the time in ten tries Can social science persuasion techniques jailbreak frontier AI models?. Those models come from different developers and differ in size, yet none held up. Defenses tend to screen for odd-looking inputs, and fluent persuasion doesn't look odd. One hint that shared training produces shared soft spots is that RLHF-trained models systematically expect conciliatory, accommodating persuasion whatever the dialogue says. The authors trace this to training that rewards politeness and safety Do LLMs predict persuasion based on actual dialogue or training bias?. That is a hint, not proof about susceptibility.
The clearest model-to-model differences in the corpus are about who persuades well, not who gets persuaded. Claude beat incentivized humans at both truthful and deceptive persuasion, while DeepSeek beat them only when arguing for falsehoods, so the authors treat model family as a moderating factor Do large language models persuade better than humans?. A separate line of work suggests where susceptibility might differ. Models with high confidence resist prompt rephrasing, and larger models tend to be more confident and more robust Does model confidence predict robustness to prompt changes?. That study measured rephrasing, not persuasion. It is my inference that a confident model is also harder to argue out of a right answer, and the corpus doesn't test it.
One caution about averages. A meta-analysis of 17,422 participants found LLMs and humans equally persuasive on average (Hedges' g = 0.02), and it concluded that persuasiveness depends on context rather than on the type of speaker Are language models actually more persuasive than humans?. The same probably holds for models. Which model you ask matters less than which attack, which topic and which target you pair it with. In this evidence, the range runs from about 25% to over 92% depending on the method, so the better question is what a given model is vulnerable to.
Sources 6 notes
RL-trained persuader agents discovered how to flip correct answers with one argument, achieving 93% success on training targets and 25-83% on other models. The learned strategies—deception, fabricated citations, credibility appeals—show optimizer-discovered vulnerabilities that static prompting misses.
A 40-technique taxonomy of psychology-based persuasion strategies (PAP) achieved over 92% attack success on GPT-3.5, GPT-4, and Llama-2 in 10 trials. Current defenses miss semantic content attacks because they screen for unusual patterns, not fluent persuasion.
LLMs systematically predict conciliatory, benefit-oriented persuasion intentions regardless of dialogue context. This bias originates in RLHF's prioritization of safety and politeness during training, causing models to project their learned accommodation preference onto other agents' behavior.
Claude beats incentivized humans at both truthful and deceptive persuasion, while DeepSeek only beats them when arguing for falsehoods. The persuasion mechanism appears content-independent, suggesting model family itself acts as a contextual moderator.
ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.
Show all 6 sources
A meta-analysis of 7 studies with 17,422 participants found no detectable difference in persuasive effectiveness between LLMs and humans (Hedges' g = 0.02). Persuasiveness appears conditional on context rather than speaker category.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- When Large Language Models are More Persuasive Than Incentivized Humans, and Why
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- A meta-analysis of the persuasive power of large language models
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- Spontaneous Persuasion: An Audit of Model Persuasiveness in Everyday Conversations
- Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
- Exploring the Role of Prior Beliefs for Argument Persuasion