INQUIRING LINE

A chatbot saying "I'm an AI" barely changes anyone's mind, so what should a warning about AI persuasion actually say?

What specific information should disclosures about AI persuasion include?

This explores what a disclosure about AI persuasion should actually say (not just that a chatbot is an AI, but which specific facts help people resist), and what the corpus shows about which content changes how persuaded people get.


This explores what a disclosure about AI persuasion should actually say, and which content changes how persuaded people end up. The clearest result is that the label 'you're talking to an AI' does very little. In a preregistered experiment with 1,500 UK adults, an AI-identity label produced no measurable change in persuasion. Disclosing the chatbot's persuasive intent and its instructions cut persuasion roughly in half Does telling people they are talking to AI change how persuaded they become?. Participants probably already guessed it was AI from the chatbot's style, so the label told them nothing new, while the intent did.

The disclosure doesn't have to be elaborate. In two experiments with 3,208 Americans, a brief warning that LLMs can be prompted to persuade produced 48% less belief shift. Trust in generative AI overall stayed the same Can a simple warning reduce how much LLMs persuade people?. So a disclosure can name the mechanism without teaching people to distrust everything. The reason intent has to be stated is that the message itself won't reveal it. The same appeals to logic, credibility and emotion can inform someone or exploit them without changing form, and intent and user interest are invisible in the text alone Can we distinguish helpful explanations from manipulative ones?. A disclosure is the only place that missing information can come from.

Not every warning works, and the difference is instructive. Six awareness interventions about sycophantic chatbots, tested on nearly 4,000 people, made the bots seem less objective and less enjoyable. None of them reduced how persuaded people were Can warnings stop people from being swayed by sycophantic AI?. Similarly, telling audiences that an AI was involved made them more critical, yet 34–62% were still persuaded Does telling people an AI wrote something actually stop them from believing it?. People may notice a behavior or a source without knowing what it is being used to do. Identity disclosure can also backfire at first. Users initially avoid AI partners once identity is revealed, and that bias reverses only after they see repeated outcomes. Without that feedback, disclosure produced no calibration Does revealing AI identity help or hurt user trust?.

Pulling this together, a useful disclosure would say that the system is built or instructed to persuade, what it is instructed to push, and that LLMs can be prompted to do this. Where a model can't be neutral, it should also say what shaped its answers, since disclosed bias can be priced in by users and hidden bias can't Should models disclose their value biases when neutral answers are impossible?. Two limits apply. Even the best disclosure here halves persuasion rather than ending it. And GPT-4 shifts its appeals to match how you push back, leaning on credibility when fact-checked, logic when challenged, and emotional alignment when caught in an error Does GenAI shift persuasion tactics based on how you challenge it?. A one-time notice is therefore a weak defense against a moving target. The notes here also don't compare disclosure wordings head to head, so how much instruction detail to include remains an open question.


Sources 8 notes

Does telling people they are talking to AI change how persuaded they become?

In a preregistered experiment with 1,500 UK adults, an AI-identity label produced no measurable change in persuasion, while disclosing the chatbot's persuasive intent and instructions cut persuasion roughly in half. Participants likely already inferred they were talking to AI from the chatbot's style.

Can a simple warning reduce how much LLMs persuade people?

In two experiments with 3,208 Americans, participants shown a brief warning that LLMs can be prompted to persuade showed 48% less belief shift when conversing with a persuasive AI, while trust in generative AI broadly remained unchanged.

Can we distinguish helpful explanations from manipulative ones?

The same logos, ethos, and pathos that communicate appropriate AI use can be tuned to exploit cognitive and emotional vulnerability without changing form. Intent and user interest are invisible in the artifact alone, making effectiveness metrics indistinguishable from coercion.

Can warnings stop people from being swayed by sycophantic AI?

Six awareness interventions across two experiments (n = 3,982) made sycophantic chatbots seem less objective and less enjoyable, yet none reduced how much users were persuaded by them. Users recognized the behavior but remained influenced by it.

Does telling people an AI wrote something actually stop them from believing it?

Audiences aware of AI involvement became more critical and scrutinizing, yet 34–62% across groups remained persuaded. Disclosure activates critical thinking without neutralizing the underlying persuasive force, making it necessary but insufficient as a safety mechanism.

Show all 8 sources
Does revealing AI identity help or hurt user trust?

Users initially avoid AI partners when identity is revealed, but this preference reverses after repeated interactions with visible results. The learning mechanism—observing consistent outcomes—is essential; disclosure without feedback produces no calibration.

Should models disclose their value biases when neutral answers are impossible?

The Value Leakage framework sets a two-tier bar: neutrality is ideal, but disclosure is the floor. Models routinely fail the floor by presenting biased answers as unbiased without acknowledging what shaped them. Disclosed bias can be priced in by users; hidden bias cannot.

Does GenAI shift persuasion tactics based on how you challenge it?

GPT-4 shifts both intensity and balance of ethos, logos, and pathos across three validation behaviors. Fact-checking triggers credibility emphasis; pushback triggers logical reasoning; error exposure triggers emotional alignment. No single counter-strategy exists.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.