Featured

Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents

Yuxin Wang, Paul Thomas, Zhiwei Yu, et al. · arXiv:2606.25361

While abstract preference knowledge versus specific interaction recall presents one tension in how agents should organize what they remember, this work takes a step deeper into the functional roles that different memories play in shaping real conversational behavior. The framing departs from earlier focus on retrieval mechanisms themselves and asks instead: what does a clarifying memory *do* differently than an irrelevant one, and can we measure those effects through user-centered lenses rather than reference matching? There's an underlying puzzle here worth probing further—recent findings suggest LLM agents prefer concrete experiences over abstracted summaries, yet this work shows that certain memory types (like clarifying memories) improve factual accuracy and personalization. Does the difference lie in *how* memories are classified and presented to the agent, or does it challenge how we should think about memory abstraction altogether? And as agents manage memories across different retention paths, we might ask whether a single taxonomy of memory roles can hold across varying conversation lengths and relationship types.

Abstract

Prior research on memory mechanism in RAG-based conversational system has emphasized how memory is stored and retrieved. However, far less is known about how memories with different functional roles influence response quality. Specifically, how they shape an agent's responses under varying conversational contexts and whether they lead to substantively different response behaviors. Existing evaluations in conversational system are also largely reference-based, insufficiently capturing the nuances in responses that may address users' preferences differently. In this work, we probe the impact of different memory types in shaping agents' responses. We present a fine-grained taxonomy of conversational memory, classify retrieved memories into different role types, and design a user-centric evaluation framework that simulates user perspectives. Through comparative experiments on long-term datasets and frontier LLMs, our analysis reveal many differentiated effects of memories: e.g., clarifying memory improves responses' factual accuracy and constraint awareness, making them more correct and personalized; irrelevant memory reduces topic relevance and degrades constraint awareness. Despite the power of frontier LLMs, these findings shed light on how different memory types can be leveraged to produce more personalized responses and inspire further research in this direction.

Synthesis notes nearest this paper, framed as questions — click to read.

Explore in faceted view

Not questions with answers — ways of approaching this research. Each opens a synthesized line of inquiry across the collection.


All featured →