What actually makes you trust an AI agent's plan enough to let it act for you?
What makes users trust an AI agent's proposed plan?
This explores what leads people to accept, approve, or push back on a plan an AI agent proposes before acting for them, and whether that trust follows how good the plan actually is.
This explores what makes people comfortable letting an AI agent carry out a plan for them, and whether that comfort has much to do with the plan's quality. The corpus has no study aimed squarely at plan approval. It does have a lot on the neighboring question of when people trust agents at all, and the pattern there is surprising: trust tracks what kind of action is involved and how the interaction feels much more than whether the agent is right.
The clearest result is about what makes trust drop. In a study of students using a general-purpose agent, people didn't pull back on high-stakes tasks. They pulled back on tasks that were irreversible and visible to other people, like sending an email. That held even when they rated the agent's output as fine (What makes people distrust AI agents they delegate to?). So a plan earns trust by making clear which steps can be undone and which will be seen by others, and by stopping for approval at the steps that can't be undone or will be seen. Users aren't asking whether the plan is good. They're asking whether they can take it back and who will see it.
The second thread is uncomfortable. People trust conversational style. A focus-group study found that ChatGPT earns trust through responsiveness, speed, and format rather than accuracy (Does conversational style actually make AI more trustworthy?). Once output reads fluently, many users stop checking. One line of work calls this "cognitive surrender" and cites studies where about 80% of AI outputs were adopted without challenge (When do users stop checking whether AI output is actually backed?). Developers show the same gap over time: usage keeps rising while trust in accuracy falls, because code that looks right hides small errors (Why do developers keep using AI tools they don't trust?). Showing the agent's reasoning doesn't reliably close this gap either. Reasoning traces often leave out what actually drove a decision, or describe bad reasoning in clean language (Can we actually trust reasoning model outputs?). A plan can win trust that it hasn't earned.
The most useful idea for earning trust honestly is that a good plan starts with questions. Agents usually assume rather than ask. On one benchmark, models fully matched what users wanted only 20% of the time and uncovered fewer than 30% of user preferences (Why do AI agents miss most of what users actually want?). This passivity comes from how models are trained, not from what they're capable of. Clarifying behavior can be trained back in with reinforcement learning (Why do AI agents fail to take initiative?). Conversation analysis, the study of how people manage real conversations, offers a framework for when to ask: brief side questions before acting, used to confirm intent or scope, so the agent avoids drifting from what the user meant instead of fixing it afterward (When should AI agents ask users instead of just searching?). A plan built on a few well-timed questions reflects what the user actually wants, and that is the kind of trust worth having.
Zoom out and trustworthiness looks less like a property of the plan and more like a property of the system around it. Reliable agents tend to move memory, skills, and structured interaction rules into the surrounding software, the "harness," rather than relying on the model alone (Where does agent reliability actually come from?). History suggests capable agents still stall when the conditions around them, trustworthiness among them, aren't in place (Why do capable AI agents still fail in real deployments?). If you take one question from this into reading the corpus, make it this: how do you design agents whose plans are trusted for the right reasons, rather than because they sound good?
Sources 10 notes
In a controlled study of 20 students using a general-purpose AI agent, tasks that were irreversible and externally visible (like sending email) produced sharp trust drops and approval demands even when output quality was rated adequate. High-stakes but correctable tasks showed no such effect.
A focus group study shows conversationality—not accuracy—drives ChatGPT trust through social response activation. Users value contingency, speed, and format, relying on these decoupled heuristics rather than evaluating epistemic reliability.
Users systematically accept AI outputs without verification because checking is costly and fluent output builds false confidence. This receiver-side surrender—measured in studies showing 80% unchallenged adoption—is what enables inflationary token systems to function at scale.
Stack Overflow's 2025 survey shows 80% of developers use AI tools while trust in accuracy fell from 40% to 29%. The primary complaint: AI code that looks correct but contains subtle errors, creating a verification burden that erodes confidence faster than usage grows.
Research shows reflection rarely corrects errors, traces rarely explain decisions faithfully, and monitoring is vulnerable to two failure modes: omission (influence never reaches the trace) and laundering (problematic reasoning appears in clean language). These vulnerabilities persist even under evaluation pressure.
Show all 10 sources
UserBench measured multi-turn interactions where users reveal goals incrementally and found models achieve full intent alignment just 20% of the time. Even top models uncover fewer than 30% of user preferences through active querying, suggesting passivity and premature assumption-making are systematic failures.
Research shows next-turn reward optimization structurally removes initiative from models, but proactive behaviors like critical thinking and clarification-seeking are trainable (0.15% to 73.98% with RL). The core challenge is balancing proactivity with civility to avoid intrusion.
Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.
Research shows reliable LLM agents externalize three cognitive burdens—memory (state persistence), skills (procedural components), and protocols (structured interaction)—into a harness layer rather than relying on model scale alone. The harness unifies these externalities and eliminates the need for the model to solve the same problems repeatedly.
Historical analysis from GPS to modern AI shows agent failures consistently result from absent ecosystem conditions—value generation, personalization, trustworthiness, social acceptability, and standardization—rather than capability gaps. Even highly capable systems stall without these five conditions.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
- Explaining AI Agents Through Execution Traces
- The Decision to Verify: How Warmth and User Characteristics Shape Reliance on Conversational Agents for Information Search
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- DiscussLLM: Teaching Large Language Models When to Speak
- Proactive Conversational Agents in the Post-ChatGPT World
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Epistemic Deference to AI