Fully Autonomous AI Agents Should Not be Developed

Paper · arXiv 2502.02649
LLM Alignment

We argue that fully autonomous AI agents should not be developed. In support of this position, we build from prior scientific literature and current product marketing to detail the ethical values at play in the use of AI agents, documenting trade-offs in potential benefits and risks. Our analysis reveals that risks to people increase with the autonomy of a system: The more control a user cedes to an AI agent, the more risks to people arise. Particularly concerning are risks associated with the most extreme form of full autonomy, where the lack of human constraints leads to severe risks impacting multiple human values.

Introduction. The sudden, rapid advancement of Large Language Model (LLM) capabilities – from writing fluent sentences to achieving increasingly high accuracy on benchmark datasets – has led AI developers and businesses alike to look towards what comes next. The tail end of 2024 saw “AI agents”, autonomous goal-directed systems, begin to be motivated and deployed as the next big advancement in AI technology.1 Many recent AI agents are constructed by integrating LLMs into larger, multi-functional systems, capable of carrying out a variety of tasks to achieve goals. A foundational premise of this emerging paradigm is that computer programs need not be constrained to actions explicitly defined by a human operator; rather, systems can autonomously combine and execute multiple tasks without direct human involvement. This transition marks a fundamental shift towards systems capable of creating context-specific plans in previously unspecified environments.

Discussion / Conclusion. The history of nuclear close calls provides a sobering lesson about the risks of ceding human control to autonomous systems.36 For example, in 1980, computer systems falsely indicated over 2,000 Soviet missiles were heading toward North America. The error triggered emergency procedures: bomber crews rushed to their stations and command posts prepared for war. Only human cross-verification between different warning systems revealed the false alarm. Similar incidents can be found throughout history. Such historical precedents are clearly linked to foreseeable benefits and risks of AI agents. We find no clear benefit of fully autonomous AI agents that can operate outside of human-defined constraints, but many foreseeable harms from ceding full human control. Looking forward, this suggests several critical directions: 1. Adoption of a spectrum of AI agent autonomy: Recognizing the nature of autonomy and how it interacts with system implementations can help to highlight opportunities and associated risks aligned with different values. Further, clear distinctions between levels of agent autonomy can aid in task delegation and prioritization for governance and development (Srikumar et al.). 2.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do multi-agent systems achieve genuine cooperation and reasoning? How do we evaluate AI systems when user perception misleads actual performance? How should human oversight be integrated with autonomous AI systems? How should conversational agents balance goal-driven initiative with user control? How do professional roles and expertise transform with AI-generated content? How does AI adoption affect human skill development and labor equality? When should tasks involve human-AI partnership versus full automation? How can humans calibrate appropriate trust in AI systems? How can language models sustain linguistic synchrony and intersubjectivity during dialogue? Why do agents confidently report success despite actually failing tasks? Why do multi-turn conversations degrade AI intent and coherence? Can AI systems develop genuine social understanding without embodiment?