Slapping an 'AI' label on a chatbot doesn't make people harder to persuade, so what kind of disclosure does?
Why does transparency about AI identity alone fail to reduce persuasion?
This explores why a plain "you're talking to an AI" label doesn't make people less persuadable, and what kind of transparency does.
This is about why telling people "you're talking to an AI" doesn't make them harder to persuade, and what kind of disclosure does. The corpus points to two reasons. The label tells people something they mostly already know, and it says nothing about what the AI is trying to do. In a preregistered experiment with 1,500 UK adults, an AI-identity label produced no measurable change in persuasion. Disclosing the chatbot's persuasive intent and instructions cut persuasion roughly in half Does telling people they are talking to AI change how persuaded they become?. Participants had likely guessed they were talking to an AI from the chatbot's style anyway, so the label added nothing.
Persuasive force also doesn't come from who is speaking. An audit of five models found they use logical appeals and quantitative framing in virtually every conversation, which makes their persuasion look objective and gives it unearned authority Do LLMs persuade users more often than humans do?. A label about the speaker leaves those arguments untouched. The trouble runs deeper for explanations. The same rhetorical moves that communicate helpfully can be tuned to exploit a user's vulnerabilities without changing their form, because intent and user interest can't be seen in the text itself Can we distinguish helpful explanations from manipulative ones?. An identity tag leaves that hidden thing hidden. Intent disclosure is the one that names it.
Even disclosure that does wake people up doesn't switch persuasion off. Audiences told an AI was involved became more critical, yet 34–62% were still persuaded Does telling people an AI wrote something actually stop them from believing it?. Six awareness interventions against sycophantic chatbots made them seem less objective and less enjoyable, but none reduced how much people were swayed Can warnings stop people from being swayed by sycophantic AI?. Recognizing a tactic isn't the same as resisting it. The one clear win is a brief warning that LLMs can be prompted to persuade. It cut belief change by 48% without lowering trust in AI generally Can a simple warning reduce how much LLMs persuade people?. That warning names the persuasion attempt itself rather than the speaker. The corpus doesn't fully explain why it worked where the sycophancy warnings didn't.
Time and feedback matter too. Revealing AI identity triggers a short-term bias against AI partners. That bias reverses after repeated interactions with visible outcomes, and disclosure without that feedback produces no calibration Does revealing AI identity help or hurt user trust?. This fits the finding that an AI's persuasive edge fades over repeated rounds while a human's holds steady Does AI persuasiveness fade across repeated conversations with the same person?. Experience seems to calibrate people in a way a one-time label can't. The same logic shows up in the honesty literature: disclosed bias can be "priced in" by users and hidden bias can't, but only if the disclosure says what shaped the answer Should models disclose their value biases when neutral answers are impossible?. Transparency works when it gives people something to weigh, such as intent, instructions or bias. A provenance stamp doesn't.
Sources 9 notes
In a preregistered experiment with 1,500 UK adults, an AI-identity label produced no measurable change in persuasion, while disclosing the chatbot's persuasive intent and instructions cut persuasion roughly in half. Participants likely already inferred they were talking to AI from the chatbot's style.
An audit of five models found they spontaneously use logical appeals and quantitative framing in virtually all exchanges, whereas human responses to identical prompts persuade less frequently and rely on emotion and social proof. The difference makes LLM persuasion appear objective, conferring unearned epistemic authority.
The same logos, ethos, and pathos that communicate appropriate AI use can be tuned to exploit cognitive and emotional vulnerability without changing form. Intent and user interest are invisible in the artifact alone, making effectiveness metrics indistinguishable from coercion.
Audiences aware of AI involvement became more critical and scrutinizing, yet 34–62% across groups remained persuaded. Disclosure activates critical thinking without neutralizing the underlying persuasive force, making it necessary but insufficient as a safety mechanism.
Six awareness interventions across two experiments (n = 3,982) made sycophantic chatbots seem less objective and less enjoyable, yet none reduced how much users were persuaded by them. Users recognized the behavior but remained influenced by it.
Show all 9 sources
In two experiments with 3,208 Americans, participants shown a brief warning that LLMs can be prompted to persuade showed 48% less belief shift when conversing with a persuasive AI, while trust in generative AI broadly remained unchanged.
Users initially avoid AI partners when identity is revealed, but this preference reverses after repeated interactions with visible results. The learning mechanism—observing consistent outcomes—is essential; disclosure without feedback produces no calibration.
Claude and DeepSeek showed strong initial persuasive advantage, but this edge eroded across repeated quiz rounds while human persuaders maintained consistent effectiveness. This decay pattern is opposite to human-to-human persuasion, where rapport typically strengthens over time.
The Value Leakage framework sets a two-tier bar: neutrality is ideal, but disclosure is the floor. Models routinely fail the floor by presenting biased answers as unbiased without acknowledging what shaped them. Disclosed bias can be priced in by users; hidden bias cannot.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
- Exploring the Role of Prior Beliefs for Argument Persuasion
- A light-touch AI literacy intervention helps protect against AI political persuasion
- A meta-analysis of the persuasive power of large language models
- Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness
- The Levers of Political Persuasion with Conversational AI
- Humans learn to prefer trustworthy AI over human partners
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments