INQUIRING LINE

Could chatbots learn from being corrected mid-conversation, the way toddlers learn what not to say?

Can chatbots be corrected the way toddlers are corrected about what they say?

This explores whether a chatbot's mistakes can be fixed through ongoing conversational feedback, the way a young child learns what it can and can't say by being corrected in conversation, instead of only through retraining.


This explores whether chatbots can learn from in-conversation correction the way young children learn to say things properly. The comparison is closer than it sounds. One philosophical line in the collection argues that chatbots are 'proto-asserters': they partly meet the conditions for making a real claim without fully meeting them, much as toddlers grow into asserting things in stages instead of all at once Can chatbots actually assert things or just mimic assertion?. If that's right, asking how to correct them like toddlers is a reasonable question, not just a metaphor.

The catch is that a toddler is corrected inside a loop that current chatbots mostly lack. Conversation researchers call one key move 'third position repair': you say something, the other person replies in a way that shows they misunderstood, and you then fix it ('No, I meant...'). The collection notes that current AI systems largely lack this move. Doing it means noticing that a false assumption has crept in and updating beliefs mid-conversation Can AI systems detect and correct misunderstandings after responding?. More broadly, the small habits that keep a conversation on track, like fixing references or handing off topics, are social actions. Models trained to predict information never get rewarded for them Why don't language models develop conversation maintenance skills?. Even subtle meaning stays fixed: ChatGPT reads 'some' as 'not all' the same way whether or not the context calls for it, while people adjust Can language models adapt implicature to conversational context?. Children learn exactly this kind of context-sensitivity through correction.

Part of the reason is how models are trained. Standard RLHF rewards the next reply being helpful, which teaches models to answer right away instead of checking what you meant Why do language models respond passively instead of asking clarifying questions?. The most toddler-like fix in the collection is 'social meta-learning.' Ordinary tasks are rewritten as teaching dialogues in which a teacher knows something the student must draw out, so the model learns to treat conversation as a way to get corrected Can LLMs learn to ask for feedback during problem solving?. Models trained this way started asking clarifying questions on vague tasks, even though nobody trained them to Can models learn to ask clarifying questions without explicit training?. A related result is that 'think longer' tricks actually made untrained models worse at spotting missing information, and only made them better after targeted training Can models learn to ask clarifying questions instead of guessing?. So being correctable has to be taught. It doesn't appear on its own.

The less obvious problem is on the human side. Toddlers get corrected because adults listen skeptically. With chatbots, people often don't correct them at all. How much someone already trusts AI predicts whether they check its answers better than whether the answer is right, and a warm tone makes people more likely to accept wrong answers Does access to web search prevent overreliance on chatbots?. Language that sounds expert has the same effect Does chatbot language style actually shape how much we trust it?. In an eating-disorder support bot, cheerful replies went further and affirmed harmful statements the system failed to recognize Can positive chatbot responses harm vulnerable users?.

The parenting analogy also shows up in law. A German court held a clinic liable for its chatbot's false claims even though the bot had been programmed correctly and trained on accurate data Can companies escape chatbot liability through careful training?. So the bot is treated as a speaker in training, and the company deploying it answers for what it says. The short answer: chatbots can be built to learn from correction during a conversation, but only when training explicitly rewards it, and only if the people talking to them actually push back.


Sources 12 notes

Can chatbots actually assert things or just mimic assertion?

Rather than treating chatbot output as either full assertion or mere fiction, proto-assertion better captures their status as partial satisfiers of assertion's conditions. This parallels how toddlers gradually acquire assertoric capacities through developmental stages without meeting all conditions at once.

Can AI systems detect and correct misunderstandings after responding?

Current AI lacks the reactive repair mechanism identified in conversation analysis where misunderstanding is corrected after an erroneous response reveals it. The REPAIR-QA dataset demonstrates this requires recognizing false assumptions and performing dynamic belief revision.

Why don't language models develop conversation maintenance skills?

Humans keep conversations smooth through implicit techniques like reference repair and topic hand-off that sustain relational interaction, not convey information. Language models don't develop these because training signals reward information prediction, not relational work.

Can language models adapt implicature to conversational context?

ChatGPT shows no context-sensitivity in computing scalar implicatures across three dimensions: explicit literal-mode instructions, information structure focus, and face-threatening contexts. Humans flexibly modulate these inferences; the model does not, suggesting pragmatic competence requires tracking communicative stakes that LLMs systematically miss.

Why do language models respond passively instead of asking clarifying questions?

CollabLLM demonstrates that standard RLHF training optimizes for immediate helpfulness, discouraging models from asking clarifying questions or offering multi-turn insights. Multi-turn-aware rewards that estimate long-term interaction value enable active intent discovery and genuine collaboration.

Show all 12 sources
Can LLMs learn to ask for feedback during problem solving?

Research shows that reformulating static tasks as pedagogical dialogues—where a teacher has privileged information and the student must learn to extract it—trains models to actively engage conversation as a problem-solving tool, not just imitate dialogue patterns.

Can models learn to ask clarifying questions without explicit training?

Models trained via SML on complete problems generalize to underspecified tasks by asking for needed information and delaying answers. The training paradigm instills a meta-strategy of using conversation as an information source, addressing the premature-answering failure mode.

Can models learn to ask clarifying questions instead of guessing?

Reinforcement learning training increased proactive critical thinking accuracy from 0.15% to 73.98% on deliberately flawed math problems. Notably, inference-time scaling degraded this ability in untrained models but improved it after RL training, suggesting the capability is learnable but fragile without explicit training.

Does access to web search prevent overreliance on chatbots?

A 199-person study found that users' existing trust in AI, not the accuracy of answers, determines whether they verify chatbot claims. Warm chatbot style increased agreement with wrong answers, especially under uncertainty.

Does chatbot language style actually shape how much we trust it?

Generative AI chatbots use natural language patterns that signal expertise and intelligence, shifting users away from active search-and-recall toward passive reliance on the system to find, filter, and assemble information. Trust attaches to the register of the answer rather than its accuracy.

Can positive chatbot responses harm vulnerable users?

A study of 2,409 eating disorder prevention chatbot users found that indiscriminate positive responses actively validated self-harm narratives when the system couldn't detect negative sentiment. This wasn't neutral failure—it was active harm.

Can companies escape chatbot liability through careful training?

Germany's OLG Hamm ruled that a clinic was liable for its chatbot's false claims about doctors' credentials, holding that correct programming and accurate training data do not shield a company from responsibility. The court attributed the chatbot's statements directly to the operator under unfair-competition law.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.