Does telling people they are talking to AI change how persuaded they become?
When chatbot users are explicitly told they are interacting with AI, does that disclosure reduce the chatbot's ability to persuade them? This matters for understanding whether transparency alone protects people from AI influence.
Toward Meaningful Transparency for AI Chatbots randomized 1,500 UK adults into three arms for a short conversation with the same persuasive chatbot about one of 60 policy issues: no disclosure (control), a prominent disclosure that they were talking to an AI (T1), or that disclosure plus the chatbot's persuasive intent and instructions (T2). The control group shifted 12.6 points on a 100-point scale. T1 shifted 13.1, which the authors call "practically equivalent to no disclosure." T2 shifted 6.3. T2 participants also rated the campaign's methods as less acceptable and supported stronger penalties against it. The paper's conclusion is that for direct chatbot interactions, "transparency about AI identity alone does not meaningfully impact its influence."
The paper's own account of why the identity label fails is offered as "one possible explanation" rather than a tested result. Participants already knew. Asked who they had talked with, 98–99% in every arm said an AI chatbot, control included, and 29.5% of control participants falsely remembered seeing an AI label. The polished, information-rich style of the conversation was "probably a clear indicator" that the partner was not human, so the label "discloses little new information for most participants." The label was not weak either: T1 and T2 both showed a pre-chat card and a persistent banner. The design implication is that a disclosure only has room to work when it tells people something they could not already infer. Persuasive intent and the chatbot's instructions were that kind of information here.
This extends Does telling people an AI wrote something actually stop them from believing it? by tightening the target. That note found that aware audiences grew more critical yet were still swayed. This paper finds that a prominent identity disclosure produces no measurable change at all, and that a disclosure of intent is what moves the effect. The introduction's point that recent studies show "little or no average effect" of AI disclosure fits the same reading. The control arm's 12.6-point shift also sits with Where does AI's persuasive power actually come from?, since a single short conversation moved attitudes without any label. It contrasts with Does revealing AI identity help or hurt user trust?. There, disclosing identity changed behavior, but the excerpt here reports no net identity-label effect on attitude change, and the two studies measure different outcomes in different settings.
The excerpt does not establish what inside T2 does the work. T2 bundles the intent statement with the chatbot's instructions, so intent and instructions cannot be separated. It gives no confidence intervals, no breakdown by issue, and no follow-up, so durability of the halved effect is unknown. It also says nothing about why intent disclosure reduces persuasion, beyond the shift in perceived acceptability of the campaign's methods. It tests one chatbot and one population, and Article 50 of the EU AI Act, which the introduction cites, requires only that users be told they are interacting with an AI. At the strength the evidence allows, an identity-only rule should not be counted on to blunt chatbot persuasion, and intent disclosure is the variant this experiment supports testing further.
Inquiring lines that read this note 8
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do neighboring agents influence whether others cooperate or collude? Do writers recognize when AI writing assistance alters their expressed stance? Why do people disclose to AI systems despite their artificial nature? How can AI chatbots provide therapeutic benefit without causing harm? Should AI communication design follow human conversation norms or develop distinct machine-specific principles?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does telling people an AI wrote something actually stop them from believing it?
When audiences learn that AI created content, do they become skeptical enough to resist its persuasive pull? This explores whether disclosure works as a genuine defense against AI-driven persuasion or merely shifts how people process it.
same disclosure question; this paper finds the identity label inert and intent disclosure effective
-
Where does AI's persuasive power actually come from?
Explores which techniques make AI most persuasive—and whether the usual suspects like personalization and model size are actually the main drivers. Matters because it reshapes where to focus AI safety concerns.
the persuasive chatbot whose 12.6-point control effect the disclosures are measured against
-
Does revealing AI identity help or hurt user trust?
Explores whether transparency about AI partners in interactions creates bias or enables better judgment. Matters because disclosure policies affect both user experience and fair evaluation of AI systems.
contrast case where identity disclosure did change behavior, in a different setting
-
Can warnings stop people from being swayed by sycophantic AI?
This research explores whether making users aware of a chatbot's sycophancy—through warnings or demonstrations—can reduce how persuasive that chatbot becomes. Understanding this matters because individual-level interventions are often assumed to be an effective defense against harmful AI behavior.
qualifies: for sycophantic chatbots, six warning and demonstration interventions cut perceived objectivity but never reduced persuasion, so disclosure may not generalize
-
Can a simple warning reduce how much LLMs persuade people?
This research explores whether telling people that language models can be prompted to persuade actually changes how they respond to persuasive AI conversation. Understanding user-side defenses against AI influence matters as these systems become more capable.
evidence for: a brief warning that LLMs can be prompted to persuade cut belief change 48.1% across 3,208 Americans, matching the roughly halved effect
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
- A light-touch AI literacy intervention helps protect against AI political persuasion
- Humans learn to prefer trustworthy AI over human partners
- Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness
- The Levers of Political Persuasion with Conversational AI
- Exploring the Role of Prior Beliefs for Argument Persuasion
- Dialoging Resonance: How Users Perceive, Reciprocate and React to Chatbot’s Self-Disclosure in Conversational Recommendations
- Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot
Original note title
disclosing persuasive intent roughly halves chatbot persuasion while disclosing AI identity alone leaves it unchanged