Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
To anticipate the socio-technical risks posed by AI agents, organizations first need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture the job-specific risks introduced by agents. To address this gap, we make three main contributions. First, we developed a multi-layer framework based on a review of the literature on AI agents. The framework models three core components and their interactions: the agents, their goals, and their environment. Second, we embedded this framework in a structured prompt and applied it to descriptions of 2,078 job tasks from the O*NET occupational database, producing 8,356 risk scenarios labeled by severity and deployment mode (automation or augmentation). We validated these scenarios with 45 workers across 10 job roles and an independent LLM judge, confirming their high plausibility and alignment with the corresponding job tasks. Finally, we extended an existing taxonomy to create a 15-category taxonomy of workplace AI agent risks that covers all our risk scenarios. Our analysis highlights four findings. First, augmentation is not inherently safe because overreliance on agents can gradually erode workers’ skills and oversight.
Introduction. AI agents are autonomous entities that perceive their environment, interact with humans or other agents, and act to achieve goals with varying degrees of independence (Sap- kota, Roumeliotis, and Karkee 2026). They are increasingly deployed to support everyday workplace tasks (Eloundou et al. 2024; Xi et al. 2025; Mohney 2025). AI agents do not operate in isolation: together with the environment they perceive, the goals they pursue, and the humans they work alongside, they form what we call an agentic AI system (AAIS) (IBM 2024). As these systems are deployed across jobs and industries, organizations need ways to identify and classify the risks they introduce (Weidinger et al. 2023). This need arises when teams red-team agents before launch (Ganguli et al. 2022), conduct impact assessments (Moss et al. 2021), or conduct risk and incident audits (Raji et al. 2020; McGregor 2021).
Discussion / Conclusion. We first discuss how our findings extend theory on risk of AI agents (§5.1), then translate them into practical guidance for workplace governance (§5.2), and finally examine the study’s limitations and directions for future research (§5.3). Make agentic decomposition explicit. Rather than treating an AI system as a black box, our framework models risk as emerging from the properties of three core components (agents, goals, environments) and from the interactions between them and humans. This distinction matters because interaction-driven risks can arise even when no single component is malfunctioning, a pattern our empirical findings confirm. This extends work focused primarily on agent–agent communication (Hammond et al. 2025) by showing that goal and environment mediation are equally important risk pathways. Further, the framework decomposes human interaction into five distinct relationship types (agent–human, human–human, human–environment, etc.) rather than treating it as a single category.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do multi-agent systems achieve genuine cooperation and reasoning? How do we evaluate AI systems when user perception misleads actual performance? How should human oversight be integrated with autonomous AI systems?- What makes some autonomy levels more valuable than others?
- Why does human-AI collaboration preserve safety compared to autonomous self-improvement?
- Can targeted human oversight work better than full autonomy or micromanagement?
- Can transparent and aligned AI reduce consciousness attribution by users?
- Which interaction design changes most effectively prevent consciousness attribution?
- Why does system-level alignment fail to address consciousness attribution directly?
- What role does user interface framing play in consciousness perception?
- What downstream claims about AI welfare follow from choosing one individuation scheme?
- Do anthropomorphic features like names drive consciousness attribution more than voice?
- What responsibility do designers bear for consciousness attribution risk?
- What measurable harms occur when users interact with AI as if it were conscious?
- Can design choices reduce harm without resolving the consciousness question?
- How does the philosophical distinction between simulation and realization affect liability?
- Can AI systems execute strategies without conscious intention behind them?
- How do anthropomimetic design features trigger System 1 cognitive traps?
- What are the three dimensions of anthropomimesis and their harms?