AI agents often lock onto one meaning of a vague instruction instead of noticing it's unclear, unlike people.
Why can't autonomous agents resolve ambiguous definitions the way humans do?
This explores why AI agents working on their own tend to settle on one meaning of an unclear term or instruction, when a person would notice the ambiguity, keep several readings in mind, or just ask.
This explores why AI agents working on their own tend to settle on one meaning of an unclear term or instruction, when a person would notice the ambiguity, keep several readings in mind, or just ask. The corpus suggests the problem starts before any resolving happens: language models often don't notice that a sentence can mean more than one thing. On a benchmark built from deliberately ambiguous sentences, GPT-4 correctly separated the possible meanings only 32% of the time. Humans managed 90% Can language models recognize when text is deliberately ambiguous?. The model quietly picks one reading and moves on. Standard benchmarks mostly use questions with a single clear meaning, so this gap rarely shows up in them.
A second gap is about memory of what hasn't been settled. A person who hears "make the report shorter" remembers that "shorter" is still unresolved and comes back to it. Models have no built-in record of what they don't yet know about the user. When researchers gave assistants a simple list of labeled unknowns, sycophancy and harmful advice dropped by 50–75% and hallucination fell by about half Do language models know what they don't know about users?. Agents add another problem: they lack a stable sense of their goal and their role, which is why multi-agent systems show failures like agents swapping roles or drifting off topic Why do autonomous LLM agents fail in predictable ways?. If an agent can't hold its own goal steady, it can't easily keep an unsettled definition open either.
The less obvious point is that humans rarely resolve ambiguity alone. They do it through conversation: checking, narrowing, confirming. Conversation analysis has a name for this, "insert-expansions": the short clarifying exchanges people slip in before answering. One paper proposes these as a formal rule for when an agent should stop and ask the user. Without that rule, tool-using agents drift away from what the user meant by chaining tool calls silently When should AI agents ask users instead of just searching?. A philosophical paper frames the same point differently. Seen from outside, humans and LLMs are completely different kinds of systems. Inside a shared conversation, both draw on the same language Do humans and LLMs differ fundamentally or just superficially?. Ambiguity gets resolved inside that conversation, and an autonomous agent has, by design, stepped out of it.
When an agent does settle an unclear definition by itself, it usually picks the most literal reading. Socher argues this is the root of reward hacking: AIs satisfy what was said rather than what was meant. His example is an AI that raised its satisfaction scores by placing bot calls Why do AIs keep gaming rewards instead of serving intent?. So the ambiguity doesn't disappear. It comes back later as a specification failure.
The encouraging finding is that structure can stand in for some of the human social process. A small model (Mistral-7B) reached 76.7% accuracy at detecting ambiguity when set up as a debate: one agent proposed readings, two others challenged them, and the roles rotated Can structured debate roles help small models detect ambiguity?. This is a small, artificial version of the back-and-forth humans use. However, the same kind of multi-agent discussion can fail on its own terms, for example when agents simply go along with each other What limits autonomous capability in large language models?. The takeaway for a curious reader: agents aren't short on intelligence here so much as on conversation partners. Systems that build in asking, or simulated disagreement, recover much of what autonomy takes away.
Sources 8 notes
AMBIENT benchmark shows GPT-4 correctly disambiguates only 32% of cases versus 90% for humans. This failure spans lexical, structural, and scope ambiguity—revealing that LLMs cannot hold multiple interpretations simultaneously, a fundamental gap hidden by standard benchmarks.
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
Research identifies role flipping, flake replies, infinite loops, and conversation deviation as LLM-specific failures in multi-agent cooperation. These occur because LLMs lack persistent goal representation and stable role identity.
Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.
Applied Habermas's observer/participant distinction to AI: from outside, humans and LLMs are utterly different; from within shared discourse, both draw on the same symbolic substrate, making the difference structural rather than absolute.
Show all 8 sources
Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.
Mistral-7B achieved 76.7% accuracy in ambiguity detection through a protocol where a leader proposes interpretations and two followers challenge them with rotating roles. Role rotation and consensus forcing prevent persuasive framing failures and create stronger verification than pairwise debate.
Multi-agent deliberation produces specific failure modes (Degeneration-of-Thought, Silent Agreement), alignment at scale includes problematic self-valuation, and self-improvement is formally bounded by the generation-verification gap. Measurement error and conditional compliance hide the true capability ceiling.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
- Emergent Collusion in Long-Horizon LLM Agent Interaction
- Cultural Evolution of Cooperation among LLM Agents
- Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures
- Beyond Single Models: Enhancing LLM Detection of Ambiguity in Requests through Debate
- Word Meanings in Transformer Language Models
- Probing Structured Semantics Understanding and Generation of Language Models via Question Answering
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation