If someone never leaves your AI assistant to search the web, does that mean they're satisfied — or just stuck?
Do contained assistant sessions without web steps indicate user satisfaction?
This explores whether an AI assistant session that never leads the person out to search or browse the web means they got what they needed, or whether 'they didn't leave' is a weak stand-in for 'they were satisfied.'
This explores whether an assistant session that never sends the user out to search or browse means they were satisfied. The short answer from this collection: probably not on its own. No paper here tests that link directly. What the collection does have points to two problems with reading a self-contained session as a happy one: the session may not be self-contained in the first place, and stopping is a weak signal of success.
Start with where these sessions sit in a person's day. A cross-surface panel study Do people use AI assistants before or after searching? found that assistant use comes *after* search and browsing much more often than before it, by 20.6 percentage points. That runs against the popular idea of the assistant as an 'answer engine' that replaces search. The same study found that assistant-only sessions are more common than search-only sessions among the same users. So a session with no web steps inside it may still be the last step of a longer journey that began on the web. The person may have arrived already informed, and the assistant was used to synthesize or finish the job. 'No web steps' describes where the session's boundary falls. It doesn't tell you whether the assistant met the need by itself.
The second problem is with satisfaction. Work on the STORM system Does user satisfaction actually measure cognitive understanding? found that people report being satisfied even while they're still confused, especially when they don't know what they're missing. The better sign of real understanding was continued engagement, not a quick sense of being done. This suggests a counterintuitive reading: a user who never goes looking for more may sometimes be someone who doesn't realize there was more to find. A tidy, contained session fits that pattern just as well as it fits a satisfied one.
Other studies in the collection show why one behavioral signal rarely captures 'it went well.' On phone agents Do phone agents succeed at all three critical tasks equally?, task success, privacy-safe completion and reuse of saved preferences turned out to be separate abilities, and ranking agents by success alone didn't predict the other two. Satisfaction also changes over time. Long-term research on personalized chatbots Does chatbot personalization build trust or expose privacy risks? shows each good interaction raises what users expect next, so the same session can satisfy someone early on and disappoint them later. Some failures also leave no visible trace. When tool-using assistants chain actions together silently When should AI agents ask users instead of just searching?, they can drift from what the user meant without the user ever stepping in. A session can look smooth and contained while answering a slightly different question than the one asked.
What you might not have expected: within this collection, an assistant session with no web steps is better read as a clue about *where a person is in their search journey* than about how satisfied they are. Testing satisfaction properly would mean pairing that behavior with something else, such as whether the person comes back, keeps exploring, or actually understood the answer. That pairing is the gap the corpus leaves open.
Sources 5 notes
A cross-surface panel study found assistant sessions come after search and browsing 20.6 percentage points more often than before, reversing the "answer engine" narrative. Assistant-only sessions are also more common than search-only sessions within the same users.
STORM shows users express satisfaction despite internal confusion, especially when unaware of knowledge gaps. Sustained engagement correlates with actual self-understanding, not immediate satisfaction ratings.
MyPhoneBench demonstrates that task success, privacy-compliant completion, and saved-preference reuse are statistically distinct capabilities with no model dominating all three. Success-only rankings do not predict privacy or preference performance.
Longitudinal research shows personalization enhances trust and anthropomorphism but also amplifies privacy concerns and escalating user expectations. One-shot studies miss these temporal dynamics—each interaction raises the baseline, making failures more disappointing.
Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do Phone-Use Agents Respect Your Privacy?
- Insert-expansions For Tool-enabled Conversational Agents
- WHEN TO ACT, WHEN TO WAIT: Modeling Structural Trajectories for Intent Triggerability in Task-Oriented Dialogue
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- The New Shape of Search: How Conversational AI Recomposes Information Seeking
- From speaking like a person to being personal: The effects of personalized, regular interactions with conversational agents
- Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
- DiscussLLM: Teaching Large Language Models When to Speak