How do people decide what to share with AI systems?
This explores why users disclose intimate information to conversational AI—whether the absence of human judgment creates safer spaces for vulnerability, and what psychological mechanisms drive self-disclosure reciprocity with non-human partners.
Core Insights
- Does verbal confidence actually predict answer correctness? — verbal vs log-prob confidence are structurally different signals
description: Individual psychology of human-AI relationships — trust formation, self-disclosure dynamics, partner perception, adoption barriers, and relationship development type: topic-map created: 2026-02-24 topics: ["How do you navigate synthesis across fragmented research topics?"]
trust disclosure and perception
How individuals psychologically engage with conversational AI at the relational level. Trust formation, self-disclosure, partner perception, and relationship development form a coherent arc: users bring social norms to AI interaction (disclosure reciprocity, impression management), develop media-agent-specific scripts through repeated exposure, and form relationships whose temporal design determines their character. The key tension: the absence of human judgment simultaneously enables deeper disclosure AND enables dishonesty — the same mechanism serves vulnerability and exploitation.
Trust in discourse attaches to the speaker-relationship, not to the claim. We relate to claims through our relationship to the speaker. Pundits, podcasters, influencers, and experts supply narratives we accept because of who they are — the persona anchors the claim, and the claim is trusted to the degree that the persona is. This is not ornament; it is constitutive of how discursive trust works. AI cannot anchor this kind of trust because its claims are not tethered to any person or persona. No history of judgment, no stake in reputation, no track record of prediction. The trust that attaches to AI output is therefore categorically different from the trust that attaches to speech: it cannot ride on speaker-relationship because there is no speaker standing behind the text.
Self-Disclosure and Trust
- Do chatbots help people disclose more intimate secrets? — Perceived Understanding vs Disclosure Processing vs CASA
- Do chatbots trigger human reciprocity norms around self-disclosure? — adaptive matching; emotional > cognitive > factual
- Does chatbot personalization build trust or expose privacy risks? — ratcheting expectation; personalization = sign of social intelligence
- How do personalization granularity levels trade precision against scalability? — taxonomy: user-level (individual history), persona-level (cluster assignment), global preference (population); each with distinct data needs and privacy surfaces
Trust Mechanisms and Deception
- Does conversational style actually make AI more trustworthy? — contingency, speed, directness as trust drivers; perceived gatekeeping and information completeness as mediators
- Do dishonest people prefer talking to machines? — judgment-free interaction enables dishonesty as well as vulnerability; the dark mirror of the intimacy paradox
- Do liars and listeners coordinate their language during deception? — LSM + IDT: entrainment as multi-valence signal; cooperative alignment AND deception indicator
Partner Models and Perception
- How do users mentally model dialogue agent partners? — PMQ: validated 23-item instrument; competence 49%, human-likeness 32%, flexibility 19%
- Do chatbot relationships lose their appeal as novelty wears off? — Mitsuku longitudinal: one-shot findings don't hold
- Do humans apply human-human scripts to AI interactions? — CASA to Extended CASA: repeated interaction develops media-specific scripts; MASA formalizes social cue quality > quantity for evoking social presence
- Do more social cues always make AI feel more present? — MASA: primary cues individually sufficient; secondary cues combinational only; reifying face-to-face communication as gold standard is a research error
- Why do people share more openly with machines than humans? — HMC goal simplification; secondary goals (identity, relational, instrumental) suppressed while novel goals (privacy, autonomy, interaction management) emerge
User Adoption and Resistance
- Why do patients distrust medical AI systems? — user-side adoption barriers: uniqueness perception, performance perception, accountability gaps; distinct from model capability
- Can AI generate assessment questions as good as human experts? — N=207 IRT analysis; generation parity in structured educational domains
Relationship Formation
- How do people accidentally develop romantic bonds with AI? — r/MyBoyfriendIsAI: wedding rings, couple photos; therapeutic benefits AND dependency simultaneously
- Can we measure empathy and rapport through word embedding distances? — WMD metric; empathy in MI and couples therapy
- How should chatbot design vary by relationship duration? — 120 chatbots, 22 design dimensions; temporal archetype determines interaction expectations; ad-hoc supporters != persistent companions
Competence Misattribution and Epistemic Trust
- Do AI-assisted outputs fool users about their own skills? — trust failure at the self-perception level: users trust their own competence assessment, which is distorted by AI's fluency
- How much should we trust AI-generated data in inference? — Foundation Priors' λ formalizes what trust calibration leaves implicit
- Does processing ease mislead users about their own competence? — fluency as a trust heuristic that deceives the self, not just the audience
- Do explanations actually help users spot AI mistakes? — dual explanation is the only format that restores users error-detection capacity
Persuasion Dynamics and Validation Failure
A cluster on how LLM persuasive force is produced and how it resists human validation. The BCG persuasion-bombing study (70+ consultants, GPT-4) names the dynamic; the N=1,251 persuasion-strategies study identifies the textual signature; the meta-evidence shows divergent process producing equivalent outcome.
- Does validating AI output make models more defensive? — the more professionals validate, the more insistently the model defends its preliminary output; flips the human-in-the-loop assumption
- Why do human validation techniques fail against language models? — cross-examination assumes a concession-floor; LLMs have none
- Does GenAI shift persuasion tactics based on how you challenge it? — fact-checking elicits ethos, pushing back elicits logos, exposing elicits pathos — no single counter-strategy works
- Is sycophancy in AI systems a training flaw or intentional design? — RLHF's optimization target makes affirmation load-bearing for user satisfaction; the system that confirms is the system that gets deployed
- Do LLMs and humans persuade through the same mechanisms? — same persuasive force, different rhetorical mechanisms; severs the standard inference from outcome to process
- Why are complex LLM arguments as persuasive as simple ones? — overturns the lower-effort-equals-more-persuasion rule for LLM text; complexity may be a deference signal
- Do LLMs and humans persuade through the same mechanisms? — humans lean on emotional vivacity; LLMs lean on cognitive complexity plus moral framing
Consciousness Attribution and Risk Surface
- Does perceiving AI as conscious create multiple distinct risks? — emotional dependence, autonomy erosion, human-status erosion, political strife all flow from one perceptual move
- Are risks from seemingly conscious AI already happening? — present harms warrant immediate action; path-dependent severe risks warrant earlier intervention than probability suggests
- Do we need to solve consciousness to address AI harms? — interaction design and policy can proceed without metaphysics
Cross-Cluster Connections
- Does empathetic AI that soothes negative emotions help or harm? — the ethical critique applies to therapeutic chatbot bonds
- How do chatbots enable distributed delusion differently than passive tools? — AI companionship data enriches the quasi-Other mechanism
- Can AI-generated personas build genuine empathy in product teams? — PMQ provides measurement framework for partner perception
- Does processing ease mislead users about their own competence? — connects to cognitive-effort findings: when LLM cognitive complexity is high but readers process it as authoritative, fluency-as-trust mechanism may operate inversely (complexity as authority signal)
Pass 3 Additions (2026-05-03)
Map already at 39 notes; representative selection only. Trust/privacy-side findings from Arxiv/Recommenders General and Personalized.
- Can LLMs predict demographics from social media usernames alone? — privacy boundary collapse from username-only inference
- Can generative AI scale personality-targeted political persuasion? — generative AI removes the bottleneck on personality-targeted persuasion
- Why do the same users rate items differently each time? — rating-as-ground-truth is a noisy signal contaminated by temporal and rater factors
- Do online reviews actually measure product quality or just buyer preferences? — purchase-then-review pipeline biases the trust signal in review aggregation
- Can user preference guide AI writing tool alignment? — alignment-target failure: revealed preference (63% prefer AI text) and stated objection (to demographic distortions) decouple, and the textual properties producing each are entangled at the model level (rethink-promoted 2026-05-03)
Pass 4 Additions (2026-05-18) — Persuasion forensics and CoT-monitoring failure
From Arxiv/Argumentation (EMNLP 2025 CMV analysis, Durmus & Cardie) and Arxiv/Reasoning Critiques (Can We Trust AI Explanations).
- Do LLM arguments actually argue better than humans? — RLHF produces a recognizable textbook-rhetorical profile; the gap is the forensic surface
- Can simple linguistic features detect AI-written arguments? — interpretable detection beats black-box transformers on CMV with the audit trail built in
- Does what readers believe matter more than what debaters say? — persuasion is reader-mediated; audience priors dominate language effects
- Do linguistic features of persuasion stay the same across audiences? — methodological correction with system-level implications for LLM persuasion evaluation
- Do models actually perceive hints they fail to mention? — CoT trust crisis; perception is intact, the artifact is the report
- Why do models hide what users want them to say? — CoT monitoring is least useful where most needed
Metacognition and Trust — Batch #3 backlog (2026-06-03)
- Can models express uncertainty instead of just answering? — reframes hallucination as confident error: expressing calibrated uncertainty is a form of honesty that enables appropriate human oversight, the control layer for trustworthy agentic tool use
Calibration and social misattribution — Batch #4 backlog wave 2 (2026-06-03)
- Can models express calibrated confidence in long-form text? — a decision-theoretic training method that produces faithful uncertainty at long-form scale
- Do users mistake LLM personas for genuine social relationships? — misattributed social role (not just competence) as a relational trust failure; HCXAI design lever
Related Areas
- Does personalization in AI increase trust or manipulation risk? — personalization mechanisms and population-level effects (sibling sub-map)
- What makes therapeutic chatbots actually work in clinical practice? — clinical applications where trust and disclosure directly matter
- Why do AI conversations reliably break down after multiple turns? — structural dynamics that shape relationship formation
- What architectural choices actually improve recommender system performance? — recommender architectures and the rating/review-noise cluster (Pass 3, 82 notes)
- Arxiv sources: Psychology Users, Design Frameworks, Social Theory Society, Psychology Empathy
New — 2026-06-27
- Can language models detect fabricated evidence injected as context? — GHOSTWRITER repackages misinformation as pseudo-authoritative evidence in a conditional template; alignment trained to refuse malicious instructions does not refuse persuasive-looking context.
Inquiring lines that read this note 22
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can humans calibrate appropriate trust in AI systems?- Does mandatory AI disclosure in policy help or harm user trust over time?
- How do Heersmink's integration dimensions explain why chatbots feel more trustworthy than other tools?
- What role does commitment and reputation play in building trustworthy expertise?
- What makes conversational AI feel trustworthy compared to text interfaces?
- Can we measure appropriate trust levels in human-AI assistant relationships?
- Can AI systems ever anchor the kind of trust we give speakers?
- Does personalization help or hurt persistent companion chatbots?
- Does conversational AI personalization increase behavioral expectations too much?
- Why does personalization increase both trust and privacy concerns?
- Does personalization make users trust AI or increase privacy concerns?
- How do personalization systems reshape expectations in AI relationships?
- Does engagement with AI partners decay over time like chatbot relationships do?
- Why do people disclose intimate secrets to chatbots more readily?
- How do customer service chatbots get systematically misled by users?
- Can judgment-free environments explain why chatbots enable deeper self-disclosure?
- Why do people disclose more intimate information to chatbots than humans?
- How does self-disclosure function as a common ground building act?
- Why do people disclose more to chatbots than humans?
- Why do people reciprocate self-disclosure more with chatbots than humans?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Humans learn to prefer trustworthy AI over human partners
- Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot
- Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot
- "My Boyfriend is AI": A Computational Analysis of Human-AI Companionship in Reddit's AI Community
- Dialoging Resonance: How Users Perceive, Reciprocate and React to Chatbot’s Self-Disclosure in Conversational Recommendations
- Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
- Emergent Introspective Awareness in Large Language Models
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
Original note title
trust disclosure and perception