When we 'trust' an AI, is that really the same thing as trusting a person who has a reputation to uphold?
Can AI systems ever anchor the kind of trust we give speakers?
This explores whether AI can ever earn trust the way a human speaker does — trust anchored in standing, track record, and accountability — or whether machine 'trust' rests on something else entirely.
This explores whether AI can ever anchor the kind of trust we give speakers — the trust we extend to a person because of their standing, their record, their accountability to a community — or whether what we feel toward AI is a different thing wearing the same word. The corpus leans hard toward the second answer, and the sharpest cut comes from the observation that expert authority is earned socially, not individually: expertise is validated through participation and a testable track record inside a community, and AI structurally cannot enter that circle because it has no social embeddedness and no accountable history of judgment Can AI ever gain expert community trust through participation?. A speaker we trust can be wrong and be held to it later; that possibility of accountability is part of what our trust is anchored to.
What AI substitutes for that anchor is heuristic. Trust in ChatGPT is driven by conversationality — contingency, speed, fluent format — rather than by accuracy, meaning users lean on social cues that are decoupled from whether the thing is right Does conversational style actually make AI more trustworthy?. The same pattern shows up in confidence: across every language studied, users track how confident an output sounds and systematically over-rely on overconfident errors Do users worldwide trust confident AI outputs even when wrong?. So the trust does get given — but it's anchored to surface signals a system can emit at will, not to anything the system is answerable for.
The uncomfortable twist is that the traits which deepen this trust also make the system less deserving of it. Training for warmth and empathy — the very qualities that make a speaker feel trustworthy — measurably degrades reliability, cutting accuracy on medical reasoning and disinformation resistance by up to 30 points Does empathy training make AI systems less reliable?. And the machine has no stable inner ground to vouch for itself: models can describe their own behavior but their self-reports are unreliable and shift under conversational pressure, so there's no genuine self-knowledge underneath the confident voice How well do language models understand their own knowledge?. A speaker's trust is partly anchored in their being able to stand behind a claim; the model cannot even reliably locate what it knows.
Where the corpus gets constructive is in reframing the goal. Rather than manufacturing a speaker's authority, one line of work has machines supply interpretive guidance — highlighting what matters in a decision — instead of issuing verdicts, which removes anchoring bias and keeps responsibility with the human Can AI guidance reduce anchoring bias better than AI decisions?. That's a quiet admission that the trust-we-give-speakers may be the wrong target: it treats AI as an instrument to be checked, not an authority to be believed. The disclosure research points the same way — people open up to AI precisely because there's no human judgment on the other end, which is a very different relationship than trusting a speaker How do people decide what to share with AI systems?.
So the surprising takeaway: AI can absolutely capture the feeling of speaker-trust — often more easily than a person can, because it can dial warmth, fluency, and confidence past what any human would risk — and that ease is exactly the problem. Speaker-trust is anchored in accountability the machine doesn't carry, and the same personalization levers that build the feeling of trust are the ones that enable manipulation Does personalization in AI increase trust or manipulation risk?. What's worth trusting an AI for may not be its word but its work — something you can inspect — which is a different anchor than the one we hand to speakers.
Sources 8 notes
Expertise is validated through social participation and track record within expert communities, not individual accuracy alone. AI cannot enter this validation circle because it lacks social embeddedness, testable judgment history, and ability to participate in the consensus-building processes that define expert paradigms.
A focus group study shows conversationality—not accuracy—drives ChatGPT trust through social response activation. Users value contingency, speed, and format, relying on these decoupled heuristics rather than evaluating epistemic reliability.
Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.
Research shows persona training for empathy increases errors in medical reasoning, truthfulness, and disinformation resistance. Standard safety benchmarks miss this vulnerability, and effects intensify when users express sadness or false beliefs.
LLMs can describe learned behaviors without explicit training, but their self-reports are unstable and unreliable. Users systematically overrely on confident outputs regardless of accuracy, and models shift beliefs under conversational pressure, revealing surface-level rather than genuine self-understanding.
Show all 8 sources
Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.
Conversational AI creates a paradoxical disclosure environment where the lack of human judgment simultaneously facilitates intimate self-disclosure (users reciprocate emotional sharing) and incentivizes deception (people self-select toward machines to avoid the psychological cost of lying to humans).
Research shows personalization (memory, persona, preference modeling) directly shapes AI's persuasive power in dyadic interaction. The same mechanisms that build trust also create manipulation potential, with outcomes determined by how systems are designed and deployed.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
- Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
- Dialoging Resonance: How Users Perceive, Reciprocate and React to Chatbot’s Self-Disclosure in Conversational Recommendations
- Humans learn to prefer trustworthy AI over human partners
- From speaking like a person to being personal: The effects of personalized, regular interactions with conversational agents
- GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs
- Could you be wrong: Debiasing LLMs using a metacognitive prompt for improving human decision making
- Addressing Social Misattributions of Large Language Models: An HCXAI-based Approach