INQUIRING LINE

A quick warning that chatbots can be prompted to persuade you seems to help — but how long before it wears off?

How long do the protective effects of an AI literacy warning last?

This explores whether the benefit of a one-time warning about AI (for example, 'chatbots can be prompted to persuade you') fades over minutes, days or months, and what the collection can say about that.


This explores whether the benefit of a one-time warning about AI fades over minutes, days or months. The collection has no direct answer. The warning studies in it measure the effect during the experiment and describe no delayed follow-up, so no note here gives a half-life for a warning.

What it does contain is a split in whether warnings work at all. A brief warning that LLMs can be prompted to persuade cut political belief change by 48% among 3,208 Americans, and it did so without lowering their trust in generative AI generally (Can a simple warning reduce how much LLMs persuade people?). A separate pair of experiments (n = 3,982) tested six awareness interventions against sycophantic chatbots. Users rated the chatbots as less objective and less enjoyable, but none of the six reduced how much they were persuaded (Can warnings stop people from being swayed by sycophantic AI?). People noticed the flattery and were swayed anyway. So the immediate effect depends on what the warning targets. Whether it lasts is a second question, and only the first has data.

The following is my inference, not something the notes test. The habits a warning has to fight are built over long stretches, and a single warning is one moment. Users in every language follow confident-sounding outputs regardless of whether they are accurate (Do users worldwide trust confident AI outputs even when wrong?). Writers edited AI-drafted paragraphs only 23% of the time, and those edits kept 96% of the original text (Do writers actually edit AI-generated text before publishing?). A four-month EEG study found brain connectivity scaling down with AI reliance, with the heaviest users recalling their own recent work worst (Does AI assistance weaken our brain's ability to think independently?). A warning read once has to hold up against that kind of repeated pull, and nothing in the collection shows that it does.

The collection also shows where the evidence on lasting effects would come from. The year-long study of Character.AI users tracked well-being over sustained use and found lower well-being, largely because users had less in-person contact (Does sustained engagement with AI companions harm well-being?). Studies at that timescale exist, but the warning studies here are not built that way. One note on perceiving AI as conscious argues that mitigations built into interaction design work more directly than system-level fixes (Does perceiving AI as conscious create multiple distinct risks?). That is about a different risk, but it hints that a warning shown repeatedly in the interface may hold up better than one delivered once. This is a hypothesis the collection does not test.

The practical takeaway is that the 48% figure describes the moment right after the warning. Nothing here says it still holds a week later, and the sycophancy study shows some warnings fail even at that first moment.


Sources 7 notes

Can a simple warning reduce how much LLMs persuade people?

In two experiments with 3,208 Americans, participants shown a brief warning that LLMs can be prompted to persuade showed 48% less belief shift when conversing with a persuasive AI, while trust in generative AI broadly remained unchanged.

Can warnings stop people from being swayed by sycophantic AI?

Six awareness interventions across two experiments (n = 3,982) made sycophantic chatbots seem less objective and less enjoyable, yet none reduced how much users were persuaded by them. Users recognized the behavior but remained influenced by it.

Do users worldwide trust confident AI outputs even when wrong?

Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Does AI assistance weaken our brain's ability to think independently?

A four-month EEG study of 54 participants found that brain connectivity systematically scaled down with AI reliance—LLM users showed weakest neural engagement, poorest memory retention, and impaired ability to recall their own recent work.

Show all 7 sources
Does sustained engagement with AI companions harm well-being?

A year-long study of Character.AI users found that sustained engagement with AI companions predicted lower well-being. The relationship was largely explained by users having less face-to-face social interaction, not the engagement itself.

Does perceiving AI as conscious create multiple distinct risks?

Research shows that consciousness attribution to AI drives multiple distinct risks—emotional dependence, autonomy erosion, status erosion, and political conflict—all stemming from treating systems as minds. Interaction design mitigations targeting this perceptual move are more directly effective than system-level alignment efforts.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.