A light-touch AI literacy intervention helps protect against AI political persuasion
Conversations with large language models (LLMs) can substantially shift beliefs and attitudes, raising concerns about manipulation using AI persuasion. Here we test whether a light-touch AI literacy intervention – a brief warning that LLMs can be prompted to persuade and may present information selectively – helps protect users. Across two experiments (total N = 3,208 Americans) in which participants conversed with an LLM instructed to shift their views about different political topics, the presence of a warning reduced belief change by roughly one-half (-48.1%, 95% CI [-59.5%, - 36.8%]) relative to the control. Importantly, the warning did not significantly reduce trust in generative AI more broadly. Light-touch literacy interventions can help protect users against AI political persuasion.
Introduction. There is increasing evidence that conversations with AI chatbots can be highly persuasive. While these conversations can be used in beneficial ways, such as debunking conspiracy theories (Costello et al., 2024) or reducing science skepticism (Hornsey et al., 2026), there is substantial concern about AI dialogues being used to persuade in a harmful manner, such as persuading voters about political issues (Hackenburg et al., 2025; Salvi et al., 2025; Argyle et al., 2025) and candidates (Lin et al., 2025; Potter et al., 2024), and eroding democratic norms (Schroeder et al., 2025). Despite these concerns, little work has developed or tested solutions to protect users from influence by conversational AI. Here we ask whether a minimal AI literacy intervention can reduce susceptibility. In two studies, we test the effect of informing participants about the potential for large language models (LLMs) to be prompted to persuade, and thus to provide biased or selective information (see Fig. 1A).
Discussion / Conclusion. Given the evidence that AI chatbots can persuade across a wide range of issues, there are widespread calls to find ways to limit these persuasive effects. Here, we present evidence that a light-touch AI literacy intervention – simply informing people that AI models may have motives to persuade or manipulate (even without providing specific information about the model’s intent) – can reduce persuasion in political settings by approximately one-half. Importantly, this intervention did not have a significant effect on trust in generative AI more broadly. This suggests that the warning is specifically conferring protection against political persuasion. While the intervention does not entirely eliminate AI’s persuasive effects, it is a proof of concept that literacy treatments can have a meaningful impact. Future work should establish how to most effectively deliver such information. It is also important to further investigate ways to reduce the influence of manipulative AI while preserving the benefit of accurate AI (making people more discerning rather than more generally skeptical; Guay et al., 2023). The lack of effect on overall trust in generative AI indicates that our treatment was targeted at least to some extent; future work should explore effects on prosocial persuasion.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does AI-generated content transformation affect public discourse quality?- Why do users override their own judgment when AI says a headline is false?
- How does social proof work differently when there is no identifiable author?
- How does AI's claim proliferation affect the quality of public discourse?
- Could false social proof from AI posts crowd out authentic influencer engagement?
- How do distorted AI versions of opinions spread through public discourse?
- How does the cultural reflex around advertising disclosure compare to AI disclosure?
- What threshold of skepticism does AI awareness actually create in audiences?
- Why does knowing something is AI-generated reduce agreement with it?
- Does AI authorship disclosure change how people respond to explanations?
- What happens when AI generates content faster than humans can verify it?
- How does AI fact-checking increase belief in false headlines users saw?
- Does mandatory AI disclosure in policy help or harm user trust over time?
- Can disclaimers alone prevent users from trusting AI outputs too heavily?
- Can belief-specific counterevidence help people resist AI persuasion attempts?
- How do ethos logos and pathos shape AI persuasion under scrutiny?
- Can content-side interventions reduce AI persuasion where disclosure labels fall short?
- How does collapsing the author-public distinction remove the audience an appeal would target?
- Can audiences learn to recognize and resist moralized AI rhetoric?
- Why do read-only formats give AI content more persuasive power?