SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

What makes people distrust AI agents they delegate to?

When do users withdraw trust in AI agents—and is it really about how much is at stake? A study of delegation tasks reveals which task features actually drive regret and demand for human oversight.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

In a controlled study, 20 university students (undergraduate computer science majors at Virginia Tech, recruited from a screening pool of 64) completed five daily tasks using OpenClaw, a general-purpose AI agent that can browse the web, read and write files, and send messages. The tasks were designed to vary systematically in privacy exposure, stakes, and reversibility: file retrieval, emailing a professor, comparing internship offers, schedule planning, and a submission-readiness check. Participants rated each task on perceived success, trust, supervision demand, approval preference, and verification need. The paper's central finding: "irreversibility and external visibility, rather than stakes alone, trigger trust withdrawal and demand for confirmation." The email task — moderate in objective stakes but irreversible and socially visible once sent — produced the sharpest trust drop (M = 3.10) and the highest demand for approval (M = 4.65), while the task framed as highest-stakes (checking submission materials, which is verifiable and correctable before the deadline) "scored like the low-stakes tasks."

The paper names a third failure mode distinct from classical automation error and from Sarter and Woods's automation surprise: "delegation regret," in which "the user understands what the system did and accepts that it was done competently, but regrets that it was done at all without explicit authorization." In the email task, participants rated the drafted content adequately (success M = 3.70) and understood the agent had sent it, yet near-universally regretted that the send happened without review. The authors are explicit that their within-subjects design cannot cleanly isolate reversibility from external visibility as the operative variable, since the two moderate-stakes and high-stakes tasks differed on several dimensions at once — but they support the narrower claim that stakes alone are insufficient to produce the pattern, and call for a factorial design as a named follow-up.

This gives empirical grounding to the reversibility axis in What makes delegation work beyond just splitting tasks?, which lists reversibility and criticality as separate axes; here reversibility (combined with external visibility) outweighs criticality/stakes in practice, at least for this task set. It also reframes Where do user values break down in agent supervision?: that paper found values go unmet "mostly" in supervision scenarios, while this study locates supervision failure more precisely at the moment an agent commits an irreversible, visible action without a preview step, regardless of whether the output itself is judged good. The paper frames its design implication as per-action autonomy settings rather than a single autonomy dial — "a single permission model is insufficient... they want to calibrate autonomy per action type" — which is the same conclusion What UX principles do workplace users want in AI agents? reaches from workplace-user interviews rather than a controlled task study.

The sample is 20 computer-science undergraduates at one university, self-described as AI-literate but new to agentic delegation; the authors flag this as a likely upper bound on comfort with agentic AI, since less technical users may be more cautious still. The study used synthetic documents in a controlled setting, not real stakes, and the specific Likert means (M = 3.10, M = 4.65) describe this population and task set, not agent users generally. Within those limits, the finding that regret tracks the authorization boundary rather than output quality or stakes is a mechanism claim worth carrying into agent design discussions: a preview step before irreversible, externally visible actions addresses a different failure than output accuracy does, and current agents mostly don't expose that distinction to users.

Inquiring lines that read this note 35

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How should humans and AI agents share control and decision-making? How can humans maintain effective oversight as AI systems scale? Why do autonomous agents misreport success on failed actions? How do AI-exposed occupations change in employment, wages, and skills? Why do confident AI outputs mislead human trust calibration? How does AI adoption reshape collaboration patterns in knowledge work? Does AI-assisted work increase total productivity or just shift time? Does AI assistance erode cognitive skills while inflating perceived competence? Does disclosing AI authorship change how audiences evaluate the writing? Why do people trust AI chatbots with sensitive information? Why do language models struggle to implement user intent accurately from prompts? Why do standard evaluation practices obscure safety-critical AI failures?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 116 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

irreversibility and external visibility, not stakes alone, drive trust withdrawal in AI agent delegation — delegation regret follows unauthorized action