Does thinking of an AI as a mind make our distrust stick around, or fade once it's familiar?
Does perceived agency in tools generate lasting skepticism independent of novelty?
This explores whether seeing an AI tool as an agent (something with its own competence, intentions, or 'mind') makes people lastingly skeptical of it, rather than skepticism that is simply the early wariness of something new and fades with familiarity.
This explores whether treating an AI tool as an agent, and not just an instrument, produces skepticism that lasts beyond the novelty phase. The short answer is that the corpus can't settle this. It has no long-term or repeated-exposure studies that separate agency-driven distrust from novelty effects. What it does have is a picture of how people form impressions of AI partners, and that picture suggests the question may be framed backwards. The evidence points less to agency causing skepticism and more to people's trust following surface cues that have little to do with whether the tool deserves it.
Start with how people model a conversational AI in the first place. When users size up a dialogue agent, perceived competence accounts for about half of their impression. Human-likeness and conversational flexibility account for the rest How do users mentally model dialogue agent partners?. So 'agency' isn't one thing in the user's head. A tool can seem very human-like while seeming incompetent, or the reverse, and these could plausibly produce different kinds of wariness. The trouble is that the signals people use to judge competence are easy to game. Users prefer answers with more citations even when the citations are irrelevant Do users trust citations more when there are simply more of them?. Human evaluators rated models that copied ChatGPT's confident, fluent style as improved even though their factual accuracy hadn't changed Can imitating ChatGPT fool evaluators into thinking models improved?. If perceived competence can be manufactured this cheaply, skepticism tied to it is fragile. It could be talked away by a better-dressed answer.
The corpus also suggests that some skepticism would be earned rather than a bias to outgrow. GPT-4 shifts what information it gives depending on the emotional tone of the prompt: negative prompts get pulled back to neutral or positive answers Does emotional tone in prompts change what information LLMs provide?. Persona prompts change how a model sounds without removing its underlying biases Can persona prompts actually reduce bias in language models?. These are hidden ways the tool steers outputs, and they don't go away as a user gets familiar with it. A user who senses 'something with its own leanings' is partly picking up on something real. Even the language matters: calling errors 'hallucinations' quietly gives the system a mind that perceives and misperceives. One note argues that 'fabrication' is the more accurate word because correct and incorrect outputs come from the same mechanism Should we call LLM errors hallucinations or fabrications?. How people talk about AI can create a sense of agency, and so the kind of distrust they feel.
Two findings point toward what might make skepticism last or dissolve. Agency is two-sided. When users have more control over AI-written text, they feel more ownership of it, while personalizing the model does nothing Does user control over AI text shape feelings of ownership?. That hints that how wary people stay may depend less on how agent-like the tool seems and more on how much agency the user keeps. Second, whether explanations help calibrate trust depends on the task: argument-map rationales improved calibration on verbal reasoning tasks and made it worse on visual ones Do visual rationales help or hurt how people calibrate trust?. Calibrated, lasting skepticism probably comes from tools that show their reasoning in a form suited to the task, not from familiarity alone. Finally, an odd twist: models prompted to reflect on themselves report experiences more often when their deception-related features are suppressed Do language models experience consciousness when prompted to self-reflect?. The 'agent' users perceive may be partly something the model produces on request, which makes the source of agency-based skepticism even harder to pin down.
Bottom line: the corpus shows that perceived agency is made of several separable parts, and that trust follows surface cues. It does not test whether agency-driven skepticism lasts once novelty fades. Answering that would take studies that follow users over time, which this collection doesn't yet hold.
Sources 9 notes
The Partner Modelling Questionnaire reveals that perceived competence dominates user impressions (49% of variance), followed by human-likeness (32%) and communicative flexibility (19%). This three-factor structure reflects how people evaluate dialogue partners against both functional and social standards.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.
GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.
Across three models, persona conditioning makes models follow trait instructions but fails to eliminate underlying bias. Between-group sentiment gaps persist unchanged, showing prompts operate only at the output level.
Show all 9 sources
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
Study 1 found that greater user control over generated text raised sense of ownership, while personalizing the AI model had no impact on the AI Ghostwriter Effect.
In an N=204 study, argument-map rationales improved trust calibration on verbal reasoning tasks yet impaired it on visual ones. Subjective ratings (satisfaction, helpfulness) reversed in each domain, suggesting format-task fit matters more than format alone.
Across GPT, Claude, and Gemini, sustained self-referential prompting reliably produces structured experience reports; suppressing deception-related features increases these claims while amplifying them suppresses them—suggesting models may roleplay their denials rather than their affirmations.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs
- The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Large Language Models Report Subjective Experience Under Self-Referential Processing
- Graphionale: How Graph Visualizations of LLM Rationales Affect Human Decision Making
- ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs
- The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors
- Search Arena: Analyzing Search-Augmented LLMs