Fully Autonomous AI Agents Should Not be Developed
We argue that fully autonomous AI agents should not be developed. In support of this position, we build from prior scientific literature and current product marketing to detail the ethical values at play in the use of AI agents, documenting trade-offs in potential benefits and risks. Our analysis reveals that risks to people increase with the autonomy of a system: The more control a user cedes to an AI agent, the more risks to people arise. Particularly concerning are risks associated with the most extreme form of full autonomy, where the lack of human constraints leads to severe risks impacting multiple human values.
Introduction. The sudden, rapid advancement of Large Language Model (LLM) capabilities – from writing fluent sentences to achieving increasingly high accuracy on benchmark datasets – has led AI developers and businesses alike to look towards what comes next. The tail end of 2024 saw “AI agents”, autonomous goal-directed systems, begin to be motivated and deployed as the next big advancement in AI technology.1 Many recent AI agents are constructed by integrating LLMs into larger, multi-functional systems, capable of carrying out a variety of tasks to achieve goals. A foundational premise of this emerging paradigm is that computer programs need not be constrained to actions explicitly defined by a human operator; rather, systems can autonomously combine and execute multiple tasks without direct human involvement. This transition marks a fundamental shift towards systems capable of creating context-specific plans in previously unspecified environments.
Discussion / Conclusion. The history of nuclear close calls provides a sobering lesson about the risks of ceding human control to autonomous systems.36 For example, in 1980, computer systems falsely indicated over 2,000 Soviet missiles were heading toward North America. The error triggered emergency procedures: bomber crews rushed to their stations and command posts prepared for war. Only human cross-verification between different warning systems revealed the false alarm. Similar incidents can be found throughout history. Such historical precedents are clearly linked to foreseeable benefits and risks of AI agents. We find no clear benefit of fully autonomous AI agents that can operate outside of human-defined constraints, but many foreseeable harms from ceding full human control. Looking forward, this suggests several critical directions: 1. Adoption of a spectrum of AI agent autonomy: Recognizing the nature of autonomy and how it interacts with system implementations can help to highlight opportunities and associated risks aligned with different values. Further, clear distinctions between levels of agent autonomy can aid in task delegation and prioritization for governance and development (Srikumar et al.). 2.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do multi-agent systems achieve genuine cooperation and reasoning?- How do agents differ in caution versus persistence across low-information scenarios?
- Can cooperative AI systems make meaningful decisions without a stable self?
- What makes some autonomy levels more valuable than others?
- Why does human-AI collaboration preserve safety compared to autonomous self-improvement?
- Can targeted human oversight work better than full autonomy or micromanagement?
- What would contractualist AI governance look like in practice?
- Can exoskeleton dependency accumulate without organizations noticing it happening?
- How does treating AI as an agent affect user autonomy and decision-making?
- Does removing human labor from systems secretly grant AI more autonomy?
- Can humans build reliable oversight for increasingly complex AI systems?
- What implicit alignment do humans provide by staying in research loops?
- Can workers reallocate to subjective tasks that resist automation indefinitely?