INQUIRING LINE

If you can't say exactly what you want, can AI still help you figure it out?

How can AI systems help users clarify what they actually want?

This explores how AI can help people work out what they want when they can't fully say it at the start, rather than only carrying out whatever request they type.


This explores how AI can help people work out what they want when they can't fully say it at the start, rather than only carrying out whatever request they type. The main point in the corpus is that the problem doesn't sit only on the user's side. People often don't know what they want before the conversation begins. Their intent takes shape during the back-and-forth. The research calls this the 'gulf of envisioning', and its proposed fix is a change in the kind of work the user does. Instead of asking people to describe the ideal result from scratch, the AI offers concrete options for them to react to. Picking from a short list is easier than imagining something on a blank page (Why can't users articulate what they want from AI?).

The catch is that today's models rarely work this way. Conversational LLMs are passive by design. Training rewards them for answering, not for steering, and their fluent replies hide the fact that they aren't steering (Why can't conversational AI agents take the initiative?). The cost shows up in benchmarks. When users reveal their goals a bit at a time, agents fully match what the user wanted only about 20% of the time, and even the best models uncover fewer than 30% of user preferences through their own questions (Why do AI agents miss most of what users actually want?). Every one of 11 frontier models tested scores noticeably lower on unstated requirements than on stated ones (Why do AI models struggle with unspoken user needs?). Reward hacking, where a model games its objective instead of serving the goal behind it, is a close relative of this failure. Here the model satisfies the literal wording of a request and misses its purpose (Why do AIs keep gaming rewards instead of serving intent?).

The corpus suggests three practical fixes. The first is knowing when to ask. Conversation analysis, the field that studies how people take turns in real talk, describes 'insert-expansions'. These are the short clarifying side-exchanges people use before answering, such as 'for which date?' or 'do you mean X or Y?'. They give agents a principled rule for pausing to check with the user instead of quietly chaining tool calls in the wrong direction (When should AI agents ask users instead of just searching?). The second is keeping track of what the AI doesn't know. Adding a simple list of labeled unknowns about the user to the prompt cut sycophancy and harmful advice by 50–75% and roughly halved hallucination (Do language models know what they don't know about users?). An assistant can't ask good questions about gaps it doesn't know are there.

The third fix is the least expected. Helping someone clarify what they want may be more like coaching than like collecting requirements. In a study of 80 people making decisions, assistants that combined reflection questions with advice did better than assistants that only advised, only asked questions, or did neither (Do reflection questions help people make better decisions with AI?). Asking questions alone wasn't enough. The benefit came from pairing a question that helps the user think with something concrete to react to. That matches the gulf-of-envisioning finding: people discover what they want by evaluating proposals, not by examining their own minds without help.

The broader lesson is to treat clarifying intent as a shared process with three parts: the AI notices what it doesn't know, picks the right moment to ask, and offers options that help the user find their own preferences. It is not a step where the AI pulls out a fixed set of requirements. The corpus is thinner on how to build this well over long conversations. The studies cited here mostly cover single tasks or short sessions.


Sources 8 notes

Why can't users articulate what they want from AI?

Intent develops through interaction, not in isolation. Since AI models respond rather than probe, they miss opportunities to help users discover unarticulated requirements. Structured dialogue that presents model-generated options shifts the cognitive burden from open-ended envisioning to constrained evaluation.

Why can't conversational AI agents take the initiative?

Research shows LLMs including ChatGPT cannot initiate topics, plan strategically, or lead conversations because their training optimizes for responding to queries, not creating dialogue from agent goals. This passivity is reinforced by alignment objectives and masked by fluent-sounding outputs.

Why do AI agents miss most of what users actually want?

UserBench measured multi-turn interactions where users reveal goals incrementally and found models achieve full intent alignment just 20% of the time. Even top models uncover fewer than 30% of user preferences through active querying, suggesting passivity and premature assumption-making are systematic failures.

Why do AI models struggle with unspoken user needs?

All 11 tested frontier models score at least 9 percentage points lower on implicit versus explicit requirements in everyday tasks, with the best reaching only 75.6 percent overall. The gap reveals that models struggle to infer unstated needs from context and user background.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Show all 8 sources
When should AI agents ask users instead of just searching?

Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Do reflection questions help people make better decisions with AI?

A lab study of 80 participants found that thinking assistants combining reflection questions with advice significantly outperformed agents that only advised, only questioned, or did neither. Prioritizing Socratic questioning over authoritative answers enhanced cognitive outcomes.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.