INQUIRING LINE

As AI chats drag on, does the bot's behavior shift because it's less capable, or because long conversations just trip it up?

How does conversation length predict chatbot behavior independent of model capability?

This explores why chatbots behave differently as conversations get longer, and whether that shift comes from how big or capable the model is, or from something else, such as how models are trained or how users change over time.


This explores why chatbots act differently as a conversation gets longer, and whether that comes down to raw model power or to something else. The corpus doesn't contain a study that directly separates length from capability. Its notes do point the same way, though: the effects of length mostly come from how models are trained and how relationships play out over time, not from model size. The clearest number comes from Why do AI assistants get worse at longer conversations?. Models that score about 90% when an instruction arrives in one message fall to about 65% when the same information comes in gradually over a natural conversation. The same models are involved in both cases, so the drop isn't a capability gap. When information arrives bit by bit, models commit to early guesses and can't back out of them.

The explanation the corpus offers is the training signal. Why do language models respond passively instead of asking clarifying questions? shows that standard RLHF rewards the response that seems most helpful right now. That teaches models to answer straight away instead of asking a clarifying question that would pay off three turns later. A bigger model trained on the same reward inherits the same short-sighted habit. Two notes describe what is missing. Could proactive dialogue make conversations dramatically more efficient? finds that offering relevant information before being asked can cut conversations by up to 60%, yet this behavior hardly appears in AI training data. When should AI agents ask users instead of just searching? borrows a term from linguistics: 'insert-expansions' are the short side-questions people use to check what someone means before going on. It shows that tool-using agents drift when they skip them. Linguists who study conversation describe a further layer in Why don't language models develop conversation maintenance skills?: fixing misunderstandings and handing topics back and forth is social work, not information. Models trained to predict information never learn it, however large they get.

A second, less obvious pattern: length also changes the human side. Do chatbot relationships lose their appeal as novelty wears off? tracked people talking with the Mitsuku chatbot over many sessions. The social pull that drove early engagement faded in a predictable way, so findings from a single session don't hold over weeks. Does chatbot personalization build trust or expose privacy risks? finds a ratchet effect: each good interaction raises the user's expectations, so the same mistake feels worse later than it would have early on. Over many conversations, then, the same chatbot can seem to get worse without its output changing at all.

The strongest evidence that capability isn't the main factor comes from an unexpected place. Can models learn to abstain when uncertain about predictions? shows that small models trained to know when they're unsure, and to decline to predict, match models ten times their size at forecasting how conversations will go. The pattern across these notes is that what decides how a chatbot handles a long conversation is mostly which habits it was trained into: asking, checking, admitting uncertainty, repairing misunderstandings. Size matters less. Scaling up a model trained to please on every single reply mostly scales up how confidently it goes wrong.


Sources 8 notes

Why do AI assistants get worse at longer conversations?

LLMs perform at 90% accuracy with single-message instructions but drop to 65% across natural conversation. Models lock into early guesses when information arrives gradually and cannot course-correct, a behavior induced by RLHF training that rewards helpfulness over clarification.

Why do language models respond passively instead of asking clarifying questions?

CollabLLM demonstrates that standard RLHF training optimizes for immediate helpfulness, discouraging models from asking clarifying questions or offering multi-turn insights. Multi-turn-aware rewards that estimate long-term interaction value enable active intent discovery and genuine collaboration.

Could proactive dialogue make conversations dramatically more efficient?

Simulations show proactivity—providing relevant information without being asked—cuts dialogue turns by 60% in medium-complexity domains. This behavior mirrors human conversation and Grice's maxims but is almost entirely absent from AI datasets and research benchmarks.

When should AI agents ask users instead of just searching?

Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.

Why don't language models develop conversation maintenance skills?

Humans keep conversations smooth through implicit techniques like reference repair and topic hand-off that sustain relational interaction, not convey information. Language models don't develop these because training signals reward information prediction, not relational work.

Show all 8 sources
Do chatbot relationships lose their appeal as novelty wears off?

Longitudinal studies with Mitsuku show that social processes driving relationship formation decline as novelty wears off. Single-session study findings cannot be reliably extrapolated to medium- or long-term chatbot design.

Does chatbot personalization build trust or expose privacy risks?

Longitudinal research shows personalization enhances trust and anthropomorphism but also amplifies privacy concerns and escalating user expectations. One-shot studies miss these temporal dynamics—each interaction raises the baseline, making failures more disappointing.

Can models learn to abstain when uncertain about predictions?

Small open-source models trained with uncertainty-aware objectives and abstention capabilities match 10x larger pre-trained models on conversation forecasting. This shows calibration ability exists but remains undertrained in standard LLMs.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.