SYNTHESIS NOTE
Topics›Argumentation›this note

Can a simple warning reduce how much LLMs persuade people?

This research explores whether telling people that language models can be prompted to persuade actually changes how they respond to persuasive AI conversation. Understanding user-side defenses against AI influence matters as these systems become more capable.

Synthesis note · 2026-09-25 · sourced from Argumentation

Across two experiments with 3,208 Americans in total, participants conversed with an LLM instructed to shift their views on political topics. Those who first saw a brief warning, that LLMs can be prompted to persuade and may present information selectively, showed roughly one-half less belief change than controls (-48.1%, 95% CI [-59.5%, -36.8%]). The warning "did not significantly reduce trust in generative AI more broadly." The authors present this as "a proof of concept that literacy treatments can have a meaningful impact," while noting that it "does not entirely eliminate AI's persuasive effects."

The intervention is deliberately thin. It tells people that models "may have motives to persuade or manipulate" without "providing specific information about the model's intent." The implied route is that a generic reminder about how these systems can be prompted changes how participants treat the conversation, though the excerpt does not test that route directly. The trust result is what the authors lean on: because broader trust did not significantly move, they read the warning as "specifically conferring protection against political persuasion" and as targeted "at least to some extent." The goal they name is making people "more discerning rather than more generally skeptical."

This sits as the demand-side counterpart to Where does AI's persuasive power actually come from?, which locates persuasive power in what developers and prompters do to the model. Here the lever is what the user is told before the conversation starts. It also contrasts with Does telling people an AI wrote something actually stop them from believing it?, where knowing an AI was involved left sway at 34-62%. What this warning discloses is different: not that an AI wrote the text, but that LLMs can be prompted to persuade and to select information. The designs, samples and outcome measures differ, so the two results are not directly comparable, but together they point toward the content of a disclosure mattering as much as whether one is made. For the risk framing in Where do frontier AI models actually pose the greatest risk today?, this is a user-side mitigation tested in one setting, political persuasion by conversation.

The excerpt leaves a lot open. It does not say how long the effect lasts, what wording or delivery format was used, how the control condition was built, or which political topics were covered. A non-significant change in trust is not proof of no change, and the authors themselves say "future work should establish how to most effectively deliver such information." They also flag as unexamined whether the warning would blunt accurate or prosocial persuasion, since the beneficial uses of AI dialogue they cite, such as debunking conspiracy theories, are exactly what a blanket warning might dampen. The implication at the strength supported: a short literacy warning is a plausible low-cost layer that roughly halves measured persuasion in this design. It is not a neutralizer, and the evidence here does not yet show it discriminates between manipulative and accurate AI.

Inquiring lines that read this note 14

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do false presuppositions and sycophancy drive persistent false beliefs in models? What factors drive AI persuasiveness and how can it be mitigated? Do language models reason like humans or mimic surface patterns? Do writers recognize when AI writing assistance alters their expressed stance? Why do people disclose to AI systems despite their artificial nature? How can AI chatbots provide therapeutic benefit without causing harm?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 71 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

a brief warning that LLMs can be prompted to persuade cuts political belief change by about half with no significant drop in trust in generative AI