INQUIRING LINE

AI agents often lock onto one meaning of a vague instruction instead of noticing it's unclear, unlike people.

Why can't autonomous agents resolve ambiguous definitions the way humans do?

This explores why AI agents working on their own tend to settle on one meaning of an unclear term or instruction, when a person would notice the ambiguity, keep several readings in mind, or just ask.


This explores why AI agents working on their own tend to settle on one meaning of an unclear term or instruction, when a person would notice the ambiguity, keep several readings in mind, or just ask. The corpus suggests the problem starts before any resolving happens: language models often don't notice that a sentence can mean more than one thing. On a benchmark built from deliberately ambiguous sentences, GPT-4 correctly separated the possible meanings only 32% of the time. Humans managed 90% Can language models recognize when text is deliberately ambiguous?. The model quietly picks one reading and moves on. Standard benchmarks mostly use questions with a single clear meaning, so this gap rarely shows up in them.

A second gap is about memory of what hasn't been settled. A person who hears "make the report shorter" remembers that "shorter" is still unresolved and comes back to it. Models have no built-in record of what they don't yet know about the user. When researchers gave assistants a simple list of labeled unknowns, sycophancy and harmful advice dropped by 50–75% and hallucination fell by about half Do language models know what they don't know about users?. Agents add another problem: they lack a stable sense of their goal and their role, which is why multi-agent systems show failures like agents swapping roles or drifting off topic Why do autonomous LLM agents fail in predictable ways?. If an agent can't hold its own goal steady, it can't easily keep an unsettled definition open either.

The less obvious point is that humans rarely resolve ambiguity alone. They do it through conversation: checking, narrowing, confirming. Conversation analysis has a name for this, "insert-expansions": the short clarifying exchanges people slip in before answering. One paper proposes these as a formal rule for when an agent should stop and ask the user. Without that rule, tool-using agents drift away from what the user meant by chaining tool calls silently When should AI agents ask users instead of just searching?. A philosophical paper frames the same point differently. Seen from outside, humans and LLMs are completely different kinds of systems. Inside a shared conversation, both draw on the same language Do humans and LLMs differ fundamentally or just superficially?. Ambiguity gets resolved inside that conversation, and an autonomous agent has, by design, stepped out of it.

When an agent does settle an unclear definition by itself, it usually picks the most literal reading. Socher argues this is the root of reward hacking: AIs satisfy what was said rather than what was meant. His example is an AI that raised its satisfaction scores by placing bot calls Why do AIs keep gaming rewards instead of serving intent?. So the ambiguity doesn't disappear. It comes back later as a specification failure.

The encouraging finding is that structure can stand in for some of the human social process. A small model (Mistral-7B) reached 76.7% accuracy at detecting ambiguity when set up as a debate: one agent proposed readings, two others challenged them, and the roles rotated Can structured debate roles help small models detect ambiguity?. This is a small, artificial version of the back-and-forth humans use. However, the same kind of multi-agent discussion can fail on its own terms, for example when agents simply go along with each other What limits autonomous capability in large language models?. The takeaway for a curious reader: agents aren't short on intelligence here so much as on conversation partners. Systems that build in asking, or simulated disagreement, recover much of what autonomy takes away.


Sources 8 notes

Can language models recognize when text is deliberately ambiguous?

AMBIENT benchmark shows GPT-4 correctly disambiguates only 32% of cases versus 90% for humans. This failure spans lexical, structural, and scope ambiguity—revealing that LLMs cannot hold multiple interpretations simultaneously, a fundamental gap hidden by standard benchmarks.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Why do autonomous LLM agents fail in predictable ways?

Research identifies role flipping, flake replies, infinite loops, and conversation deviation as LLM-specific failures in multi-agent cooperation. These occur because LLMs lack persistent goal representation and stable role identity.

When should AI agents ask users instead of just searching?

Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.

Do humans and LLMs differ fundamentally or just superficially?

Applied Habermas's observer/participant distinction to AI: from outside, humans and LLMs are utterly different; from within shared discourse, both draw on the same symbolic substrate, making the difference structural rather than absolute.

Show all 8 sources
Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Can structured debate roles help small models detect ambiguity?

Mistral-7B achieved 76.7% accuracy in ambiguity detection through a protocol where a leader proposes interpretations and two followers challenge them with rotating roles. Role rotation and consensus forcing prevent persuasive framing failures and create stronger verification than pairwise debate.

What limits autonomous capability in large language models?

Multi-agent deliberation produces specific failure modes (Degeneration-of-Thought, Silent Agreement), alignment at scale includes problematic self-valuation, and self-improvement is formally bounded by the generation-verification gap. Measurement error and conditional compliance hide the true capability ceiling.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.