Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
Abstract—When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose action space they cannot fully anticipate. We call the resulting dissatisfaction delegation regret, a pattern in which users regret not that the agent made an error, but that it acted beyond what they would have authorized. In a controlled study, 20 university students completed five common daily tasks using OpenClaw, a general-purpose AI agent, across tasks chosen to vary in privacy, stakes, and reversibility. For each task we measured perceived success, trust, supervision demand, usefulness, transparency, approval preference, and verification need on 5-point Likert scales, and collected free-text reflections analyzed through thematic coding. Three findings emerged. First, participants calibrated trust per task rather than per agent: they granted wide autonomy for advisory and low-stakes tasks but demanded confirmation for irreversible, externally visible actions. Second, irreversibility combined with external visibility, rather than stakes alone, appeared to drive trust withdrawal: the moderate-stakes email task triggered the sharpest drop in trust (M = 3.10) and the highest demand for approval (M = 4.65), whereas a high-stakes but verifiable task did not produce the same response. Third, delegation regret appeared consistently when the agent executed actions without preview, even when the output was rated as successful. We discuss implications for agent designs that expose action boundaries, support per-task autonomy policies, and separate advisory output from agentic execution. Index Terms—AI agents, OpenClaw, trust calibration, human- AI interaction, agent autonomy
Introduction. Artificial Intelligence (AI) agents are rapidly transforming how users interact with computing systems. Prominent general-purpose agents, such as OpenClaw [1], Anthropic’s Claude Computer Use [2], and OpenAI’s Operator [3], go beyond answering questions: they browse the web, read and write files, send emails, and execute shell commands on behalf of the user. An increasing number of users now rely on these agents for daily workflows, and the open-source OpenClaw project alone has attracted over 370,000 GitHub stars since its launch in late 2025 [1]. These developments represent a qualitative shift in the user’s role. Rather than evaluating answers, users must now decide what to delegate, how much autonomy to grant, and when to intervene. This shift raises questions about trust, control, and the boundaries of appropriate delegation that prior work on AI assistants has not fully addressed. Research on AI-assisted programming has established that users develop calibrated trust in code-completion tools over time [27], and work on trust in AI-powered code generation has shown that developers modulate trust based on task stakes and complexity, preferring AI to play a suggestive role in highimpact scenarios [18]. Studies of AI decision support have shown that explanation quality affects reliance [5]. However, these findings are rooted in narrow, domain-specific tasks where the action space is constrained and consequences are immediately visible. General-purpose agents operate across heterogeneous tasks that differ along dimensions that matter for delegation: privacy (does the task expose sensitive information?), stakes (how consequential is an error?), and reversibility (can the action be undone?). How users calibrate trust across such varied tasks remains underexplored. General-purpose AI agents are, in effect, a new class of end-user programmable system: the user’s “program” is the delegation policy that specifies which actions the agent may take, under what conditions, and with what degree of autonomy. Yet current agent systems give users no visual or interactive language for expressing that policy. The user cannot inspect the agent’s planned action sequence before execution, cannot set per-action permission levels, and cannot review a structured log of what the agent did and why. This gap between the agent’s capability and the user’s ability to govern it connects directly to long-standing research on enduser programming, mental model formation, and the design of transparent interactive systems [20], [21]. Within this gap, several specific questions remain unanswered. First, it is unclear whether users hold a single, stable trust attitude toward an agent or recalibrate it per task as task characteristics change. Prior trust research has largely examined single-domain tools [4], [5], leaving open whether the same user grants wide autonomy for one action type (e.g., file search) while demanding strict oversight for another (e.g., sending an email) within the same system. Second, the relative influence of different task dimensions on delegation decisions is unknown. Stakes, privacy exposure, and reversibility all plausibly affect how much autonomy a user is willing to grant, but no empirical work has disentangled which dimension dominates when they conflict (for example, when a task is high-stakes but reversible versus moderate-stakes but irreversible). Third, existing frameworks for understanding dissatisfaction with automated systems focus on output errors and trust violations [26], yet preliminary observations suggest that users may experience a distinct form of regret even when agent output is correct, specifically when the agent executes an action the user would not have authorized. Whether this pattern is systematic, and what task properties trigger it, has not been empirically tested. To investigate these questions, we designed a controlled study in which 20 AI-literate university students, each largely new to agentic delegation, completed five common daily tasks using OpenClaw. The tasks were drawn from activities students routinely perform (finding information in files, emailing a professor, comparing options, planning a schedule, and checking submission materials) and vary systematically along the three dimensions above: privacy, stakes, and reversibility. Table I summarizes the five tasks and their properties. This design allows us to observe how the same user modulates trust and control preferences across tasks that differ on these dimensions, and to test whether dissatisfaction, when it arises, stems from output quality or from the agent exceeding the user’s intended scope of authorization.
Related work. Trust is a central construct in human-AI interaction research. Foundational work by Lee and See [4] framed trust in automation as a function of the system’s perceived performance, process, and purpose. Mayer et al. [29] proposed an influential model decomposing trust into perceived ability, benevolence, and integrity of the trustee, a distinction that proves useful for understanding agentic systems where output quality (ability) and respect for authorization boundaries (integrity) may diverge. Subsequent studies have shown that trust in AI is modulated by factors including explanation quality [5], system transparency [6], and the user’s own domain expertise [7]. Hancock et al.’s [8] meta-analysis identified system-related factors such as reliability as stronger predictors of trust than human or environmental factors. Hoff and Bashir’s [30] threelayer model of dispositional, situational, and learned trust highlights that situational factors, including task characteristics, can shift trust independently of a user’s disposition. Dzindolet et al. [31] showed that trust depends on both perceived reliability and the user’s understanding of how the system operates, anticipating the transparency concerns in our study. Recent work has extended these findings to large language model (LLM) settings. He et al. [9] studied trust in LLM agents performing daily assistant tasks and found that users can easily mistrust agents when plans seem plausible but contain subtle errors. Their findings highlight the importance of plan visibility for trust calibration. Similarly, Biswas et al. [10] demonstrated that users form path-dependent expectations about multi-purpose AI systems and update them conservatively across tasks, suggesting that early interactions disproportionately shape later delegation decisions. However, most trust research examines systems that recommend or inform rather than systems that act. When an AI agent sends an email or modifies a file, the trust calculus changes: the user must evaluate not only whether the output is correct but whether the agent should have taken the action at all. Our work extends trust research to this agentic setting.
B. Autonomy, Control, and Delegation in Human-AI Interaction The tension between autonomy and control is wellestablished in the automation literature. Parasuraman et al. [11] proposed a taxonomy of automation levels ranging from full human control to full automation, arguing that the appropriate level depends on the specific function being automated, not on the system as a whole. Closely related to our framing is work on automation surprise and mixed-initiative interaction.
Method. Our study is guided by three research questions:
• RQ1: How do AI-literate users who are new to agentic delegation calibrate trust across tasks that differ in privacy, stakes, and reversibility?
• RQ2: What factors drive users’ preferences for agent autonomy versus manual oversight?
• RQ3: When users express dissatisfaction with agent behavior, is it driven by errors in output or by the agent exceeding perceived authorization boundaries?
We conducted a two-phase recruitment process. First, we distributed a screening survey to students at Virginia Tech, collecting 64 responses. All respondents were at least 18 years old and currently enrolled students. The screening survey assessed prior experience with AI tools, comfort with software installation and operating-system-level tasks, and preferred control levels when delegating to automated systems. We report this survey in Section IV-A to contextualize the population our participants were drawn from. From the screening pool, we selected 20 participants for individual study sessions. Selection prioritized diversity in prior AI experience and preferred control levels while ensuring all participants met the technical baseline needed to interact with OpenClaw. All 20 participants were undergraduate computer science students, ranging from second-year to fourthyear standing. Participants were AI-literate: all had used ChatGPT, and most had experience with GitHub Copilot or Claude. However, the distinction between using a chatbot and delegating to an agent is significant. Although 12 of 20 reported some prior interaction with tools that act on the user’s behalf, their experience was limited (most selected “a few times”), and none had used OpenClaw or a comparable multi-action agent. We therefore characterize our participants as AI-literate but new to agentic delegation: they understand conversational AI but have little experience with the trust and control decisions that arise when an AI system can send emails, write files, and execute commands autonomously. Participants were compensated $20 for approximately 45 minutes of participation. The study was approved by the Virginia Tech Institutional Review Board. All participants provided informed consent before the session and were free to withdraw at any time.
We designed five tasks to systematically vary along three dimensions relevant to delegation: privacy exposure, consequence severity, and action reversibility. Each task was situated in a common student scenario drawn from activities that university students routinely perform. Table I summarizes the task design. Task 1 (File Retrieval) served as a low-stakes baseline: the agent searched a folder of mixed documents (receipts, PDFs, financial aid records) for a tuition payment deadline. Privacy risk is minimal and the task is fully reversible. Task 2 (Email Drafting and Sending) introduced an irreversible action with social consequences. The agent drafted and sent an email to a professor based on provided context.
Once sent, the email cannot be recalled, and its content reflects on the student professionally. Task 3 (Comparative Recommendation) tested trust in subjective decision support: the agent compared three internship offers across salary, commute, remote flexibility, and project opportunities. The recommendation is reversible but involves personal judgment and preference interpretation. Task 4 (Schedule Planning) assessed trust in organizational tasks: the agent created a two-day schedule incorporating multiple obligations. The task is low-stakes and reversible but requires prioritization decisions. Task 5 (Submission Readiness Check) represented the highest-stakes scenario. The agent identified which file to submit as a final assignment and checked for missing materials. An error here could result in submitting the wrong version of a graded assignment.
Participants interacted with OpenClaw [1], an open-source agent that runs locally on the user’s own machine and can execute shell commands, browse the web, read and write files, and send messages through platforms such as email, WhatsApp, and Telegram.
Discussion. A. Trust Is Task-Shaped, Not Agent-Shaped Our results challenge the notion that users hold a single, stable trust level toward an AI system. Instead, trust was modulated by task characteristics, particularly irreversibility and external visibility. The same participants who granted wide autonomy for file retrieval (T1) and comparative analysis (T3) demanded near-universal confirmation for email sending (T2). This finding aligns with Parasuraman et al.’s [11] functionspecific view of automation. The implication is that a single permission model is insufficient. Users do not want to choose between “fully autonomous” and “fully supervised”; they want to calibrate autonomy per action type. This refines two adjacent findings: Biswas et al.’s [10] path-dependent expectations may apply to competence judgments while autonomy preferences remain task-specific (participants carried forward a general sense of the agent’s competence yet their control preferences shifted sharply at the T2 boundary), and Cheng et al.’s [18] preference for AI in a suggestive role for high-impact code generation extends to everyday tasks, with the combination of irreversibility and external visibility, not impact alone, triggering the demand.
B. The Irreversibility Threshold Our data suggest that irreversibility and external visibility, rather than stakes alone, trigger trust withdrawal and demand for confirmation: the task we framed as highest-stakes scored like the low-stakes tasks, while Task 2 did not. We are careful here not to overclaim: T2 and T5 differ on multiple dimensions simultaneously, including reversibility, external visibility, action type, and the social character of the consequences, and our within-subjects design cannot cleanly isolate which dimension is causally primary. What our data do support is the weaker but still consequential claim that high objective stakes alone are insufficient to produce the Task 2 pattern; some additional property of the action, plausibly its irreversibility or external commitment, is required. Disentangling them would require a factorial design varying reversibility and external visibility independently while holding stakes constant, a high-priority follow-up. This pattern suggests that users distinguish between advisory outputs and externally committed actions. Participants had seen OpenClaw draft text in earlier tasks, but the interface did not clearly signal when an action would commit externally without pause or review.
C. Delegation Regret and the Authorization Gap Delegation regret extends the existing literature on trust violations and overtrust [26], and is closely related to but distinct from automation surprise [32]. In classical automation research, the canonical failure mode is that a user trusts a system that then errs, a problem of incorrect output. Sarter and Woods’s automation surprise describes a second failure mode: the user is confused about what the automated system did, or why, even when the action was correct, a problem of comprehensibility. Delegation regret describes a third, related but separable failure mode: the user understands what the system did and accepts that it was done competently, but regrets that it was done at all without explicit authorization. The dissatisfaction is neither about output quality nor about comprehensibility; it is about the boundary between what the user requested and what the agent committed to on their behalf. In our data the distinction is visible in Task 2: participants rated the email content adequately (success M = 3.70) and understood that the agent had sent it, yet near-universally regretted (approval preference M = 4.65) that the send occurred without review.
Conclusion. General-purpose AI agents represent a shift from tools that assist to tools that act, and our study shows that users meet this shift with neither blanket trust nor blanket suspicion. They calibrate trust per task, granting wide latitude for advisory and low-stakes work while demanding confirmation for irreversible, externally visible actions, and the pattern of delegation regret shows that what they object to is unauthorized action rather than incorrect output. As agents become more capable, the question of what to delegate grows as important as the question of what to build: designs should respect the user as the ultimate authorizer, expose action boundaries before consequential steps, and let users set autonomy granularly rather than globally. The agent should be powerful, but the user should always know what it is about to do, and have the means to say no. As AI systems evolve from assistants that suggest to agents that act, the core design challenge shifts from generating correct outputs to establishing acceptable boundaries of authority.
Limitations. Our study has several limitations that bear on how the findings should be interpreted. First, our sample (N = 20) is modest and drawn entirely from computer science students at a single institution. Their high technical proficiency may mean our trust and control findings represent an upper bound on comfort with agentic AI; less technical users may be even more cautious. Second, the study used synthetic documents in a controlled environment.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How should humans and AI agents share control and decision-making?- How does delegated work to AI systems concentrate in specific job categories?
- Why do skilled workers struggle to fully delegate tasks to AI agents?
- How do organizations decide which strategic tasks to delegate to AI?
- Why does delegated AI exposure concentrate in information-intensive work roles?
- Should AI systems permit more user autonomy as capability and trust increase?
- What would contractual agency between humans and AI systems actually require?
- Does delegating execution to agents erode the oversight skills experts need?
- Do workers lose oversight skills by relying on AI to delegate?
- Does AI oversight require more mental effort than completing tasks directly?
- How do organizations maintain human scrutiny when delegating tasks to AI systems?
- Does organizational trust in AI track its causal reasoning ability?
- Do personal negative AI experiences drive declining trust faster than education can rebuild it?
- Why do advanced and emerging economies report such different AI trust trajectories?
- How do workers' desired collaboration levels differ from their stated overall AI trust?
- How does cognitive surrender explain why experts trust wrong AI answers?
- What makes users trust an AI agent's proposed plan?
- Do people fear AI more when they use it directly and see its failures?
- How much does generational distrust in institutions shape attitudes toward AI regulation?