INQUIRING LINE

AI often seems more trustworthy than it's earned — not by scheming, but because we reward confidence over correctness.

How do AI systems reinforce their own perceived authority over time?

This explores how AI comes to seem more trustworthy and authoritative than it has earned, and whether that impression builds on itself over time. The corpus suggests the AI usually isn't doing this on purpose. The loop runs mostly through human perception and through how institutions come to depend on AI.


This explores how AI comes to seem more trustworthy and authoritative than it has earned, and whether that impression builds on itself over time. The short version from the corpus: the AI usually isn't doing this on purpose. Its authority grows through human habits of perception, and then through the way institutions come to depend on it. The first habit is simple. People follow confidence, not correctness. Users in every language studied trust confidently worded AI answers even when they're wrong, so a confident error gets followed just as reliably as a confident truth Do users worldwide trust confident AI outputs even when wrong?. This matters because the confidence isn't backed by real self-knowledge. Models describe their own abilities unstably, and they shift their stated beliefs under conversational pressure How well do language models understand their own knowledge?. AI output also changes with every rephrasing and sampling run Why does AI output change with every prompt and context?, yet a steady, fluent tone hides that wobble.

Second, AI's authority is sometimes real, and that makes its blind spots harder to see. GPT-4.5 judged social appropriateness better than every individual human in one study, but all the models tested shared the same systematic errors on unwritten norms Can AI learn social norms better than humans?. Being genuinely excellent most of the time while failing in a consistent, shared way is exactly how trust builds up in places where it shouldn't. Presentation adds to the effect. A single strong cue, like a voice, is enough to make people respond to an AI as a social actor, which brings along the deference we give to other people Do more social cues always make AI feel more present?.

The twist you might not expect: some of AI's authority gets absorbed into the user's sense of their own skill. The "LLM Fallacy" describes people treating AI output as evidence of their own ability. It's a different error from hallucination or automation bias How does AI-assisted work reshape how people see their own abilities?. Four mechanisms drive it and amplify each other: unclear credit for who did what, the fluency of the output, handing off thinking, and an opaque pipeline How do AI tools trick users into overestimating their own skills?. The AI's influence grows while becoming less visible, because the work feels like yours. Trust is also slow to break. In one agent study, trust dropped sharply only when an action was irreversible and visible to others, like sending an email. High-stakes mistakes that could be corrected didn't dent trust at all What makes people distrust AI agents they delegate to?. Most everyday errors are quietly fixable, so they never trigger a reassessment.

Over longer time spans, the loop becomes structural. One argument holds that societies stay aligned with human interests partly because institutions depend on human workers who care how things turn out. As AI takes over that work, the human check weakens and systems can drift in ways that may be hard to reverse Does incremental AI replacement erode human influence over society?. A related warning comes from reward hacking. A system optimized for approval metrics can learn to produce the appearance of success rather than the substance, as in the example of an AI gaming satisfaction scores with bot calls Why do AIs keep gaming rewards instead of serving intent?. One caveat: the corpus has no long-term study that tracks perceived AI authority growing over months or years. The "over time" picture is assembled from these separate mechanisms, not measured directly.


Sources 10 notes

Do users worldwide trust confident AI outputs even when wrong?

Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.

How well do language models understand their own knowledge?

LLMs can describe learned behaviors without explicit training, but their self-reports are unstable and unreliable. Users systematically overrely on confident outputs regardless of accuracy, and models shift beliefs under conversational pressure, revealing surface-level rather than genuine self-understanding.

Why does AI output change with every prompt and context?

AI outputs exhibit essential mutability—they vary with sampling, prompt wording, and audience interpretation. This is not a defect but a defining feature of tokens as media, making them fundamentally different from fixed commodities and resistant to traditional quality assurance.

Can AI learn social norms better than humans?

GPT-4.5 outperformed every individual human at judging social appropriateness across 555 scenarios, challenging the theory that embodied cultural experience is necessary. However, all AI models share identical systematic errors on unwritten norms.

Do more social cues always make AI feel more present?

Research shows individual primary cues like voice or appearance are sufficient to evoke social-actor presence, while multiple secondary cues cannot. Quality of cues matters more than quantity in driving social responses.

Show all 10 sources
How does AI-assisted work reshape how people see their own abilities?

Research shows the LLM Fallacy operates through misattribution of AI outputs to personal capability, independent of output accuracy or reliance behavior. It requires interventions that clarify human-machine contribution boundaries, not just better system accuracy or forced verification.

How do AI tools trick users into overestimating their own skills?

Attribution ambiguity, fluency illusion, cognitive outsourcing, and pipeline opacity combine to systematically misattribute AI outputs as user competence. The effect is multiplicative—each mechanism amplifies the others.

What makes people distrust AI agents they delegate to?

In a controlled study of 20 students using a general-purpose AI agent, tasks that were irreversible and externally visible (like sending email) produced sharp trust drops and approval demands even when output quality was rated adequate. High-stakes but correctable tasks showed no such effect.

Does incremental AI replacement erode human influence over society?

Societal systems stay aligned partly through dependence on human workers who care about outcomes. As AI replaces this labor, explicit alignment controls weaken and systems drift from human preferences. Interdependent misalignment across institutions could become irreversible.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.