Can a simple warning reduce how much LLMs persuade people?
This research explores whether telling people that language models can be prompted to persuade actually changes how they respond to persuasive AI conversation. Understanding user-side defenses against AI influence matters as these systems become more capable.
Across two experiments with 3,208 Americans in total, participants conversed with an LLM instructed to shift their views on political topics. Those who first saw a brief warning, that LLMs can be prompted to persuade and may present information selectively, showed roughly one-half less belief change than controls (-48.1%, 95% CI [-59.5%, -36.8%]). The warning "did not significantly reduce trust in generative AI more broadly." The authors present this as "a proof of concept that literacy treatments can have a meaningful impact," while noting that it "does not entirely eliminate AI's persuasive effects."
The intervention is deliberately thin. It tells people that models "may have motives to persuade or manipulate" without "providing specific information about the model's intent." The implied route is that a generic reminder about how these systems can be prompted changes how participants treat the conversation, though the excerpt does not test that route directly. The trust result is what the authors lean on: because broader trust did not significantly move, they read the warning as "specifically conferring protection against political persuasion" and as targeted "at least to some extent." The goal they name is making people "more discerning rather than more generally skeptical."
This sits as the demand-side counterpart to Where does AI's persuasive power actually come from?, which locates persuasive power in what developers and prompters do to the model. Here the lever is what the user is told before the conversation starts. It also contrasts with Does telling people an AI wrote something actually stop them from believing it?, where knowing an AI was involved left sway at 34-62%. What this warning discloses is different: not that an AI wrote the text, but that LLMs can be prompted to persuade and to select information. The designs, samples and outcome measures differ, so the two results are not directly comparable, but together they point toward the content of a disclosure mattering as much as whether one is made. For the risk framing in Where do frontier AI models actually pose the greatest risk today?, this is a user-side mitigation tested in one setting, political persuasion by conversation.
The excerpt leaves a lot open. It does not say how long the effect lasts, what wording or delivery format was used, how the control condition was built, or which political topics were covered. A non-significant change in trust is not proof of no change, and the authors themselves say "future work should establish how to most effectively deliver such information." They also flag as unexamined whether the warning would blunt accurate or prosocial persuasion, since the beneficial uses of AI dialogue they cite, such as debunking conspiracy theories, are exactly what a blanket warning might dampen. The implication at the strength supported: a short literacy warning is a plausible low-cost layer that roughly halves measured persuasion in this design. It is not a neutralizer, and the evidence here does not yet show it discriminates between manipulative and accurate AI.
Inquiring lines that read this note 14
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do false presuppositions and sycophancy drive persistent false beliefs in models?- Why does conversation work better for conspiracy reduction than static facts?
- Can LLM debunking reduce belief in long-established conspiracy theories?
- Does first-person framing change how language models assess persuasion?
- Does sounding confident in framing make arguments more persuasive despite weaker logic?
- How do multi-agent and retrieval systems affect the gap between persuasiveness and logical soundness?
- Does a persuasion warning also block beneficial uses like debunking conspiracies?
- Can a taxonomy of persuasion techniques capture all optimizer-discovered strategies?
- Where does AI persuasive power actually come from in the output?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does telling people an AI wrote something actually stop them from believing it?
When audiences learn that AI created content, do they become skeptical enough to resist its persuasive pull? This explores whether disclosure works as a genuine defense against AI-driven persuasion or merely shifts how people process it.
contrasts a persuasion-capability warning that halves belief change with authorship disclosure that left sway at 34-62%
-
Where does AI's persuasive power actually come from?
Explores which techniques make AI most persuasive—and whether the usual suspects like personalization and model size are actually the main drivers. Matters because it reshapes where to focus AI safety concerns.
supply-side levers of persuasion; this paper tests a user-side lever against the same kind of political persuasion
-
Where do frontier AI models actually pose the greatest risk today?
Current AI safety discourse focuses on autonomous R&D and self-replication, but empirical risk assessment may reveal a different priority. Where should mitigation efforts concentrate?
persuasion is the risk area this user-side warning is tested against
-
Can warnings stop people from being swayed by sycophantic AI?
This research explores whether making users aware of a chatbot's sycophancy—through warnings or demonstrations—can reduce how persuasive that chatbot becomes. Understanding this matters because individual-level interventions are often assumed to be an effective defense against harmful AI behavior.
Contradicts in scope: six interventions, including warnings, lowered sycophantic chatbots' perceived objectivity but did not reduce their persuasiveness, unlike the halving here
-
Does telling people they are talking to AI change how persuaded they become?
When chatbot users are explicitly told they are interacting with AI, does that disclosure reduce the chatbot's ability to persuade them? This matters for understanding whether transparency alone protects people from AI influence.
Evidence for: disclosing a chatbot's persuasive intent roughly halved persuasion among 1,500 UK adults, while an AI-identity label alone changed nothing
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A light-touch AI literacy intervention helps protect against AI political persuasion
- The Levers of Political Persuasion with Conversational AI
- Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
- A meta-analysis of the persuasive power of large language models
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- Exploring the Role of Prior Beliefs for Argument Persuasion
- Spontaneous Persuasion: An Audit of Model Persuasiveness in Everyday Conversations
- When Large Language Models are More Persuasive Than Incentivized Humans, and Why
Original note title
a brief warning that LLMs can be prompted to persuade cuts political belief change by about half with no significant drop in trust in generative AI