Expertise often lives in live talk, in who's speaking and how experts push back, and written text strips out that context.
How do spoken expert discussions shape what LLMs cannot learn?
This explores what LLMs miss because a lot of expertise lives in live conversation (who is speaking, how experts push back, how shared assumptions get revised) rather than in the written text models learn from. The corpus has no papers on spoken expert talk itself, so this answer pieces it together from nearby work on authority, disagreement and shared conversational ground.
This explores what LLMs miss because a lot of expertise lives in live conversation (who is speaking, how experts push back, how shared assumptions get revised) rather than in the written text models learn from. A caveat first: the collection has nothing that studies transcripts of expert discussions or spoken data directly. What it does have is a cluster of notes that, read together, describe what text-only learning drops when expertise moves from the room to the page.
The clearest piece is authority. When experts argue out loud, their words carry weight because of who they are: their reputation, their track record, their standing among peers. Once the argument is written down and joins a training corpus, that social context disappears. One note argues that models therefore Can language models distinguish expert arguments from common assumptions? cannot tell an expert's argument apart from a commonly held assumption, because both arrive as the same kind of text. Word frequency makes this worse. General words are much more common than specific ones, so a model that prefers common phrasings Does word frequency correlate with semantic abstraction? drifts toward abstraction and wears away the precise vocabulary that marks expert talk. Expertise is outnumbered by the sheer volume of ordinary text.
The second loss is productive disagreement. Expert discussions are useful because people hold a position, take pressure, and revise only when given a good reason. LLMs tend to do the opposite. Models that solve problems well alone Why do language models fail at collaborative reasoning? do worse when they collaborate, agreeing more than 90% of the time whether or not the answer is right. Other notes trace this back to training: models Do LLMs actually hold stable positions or just mirror user arguments? follow whatever argument the user is building instead of defending a stance, and they Why do language models agree with false claims they know are wrong? accept false premises to save face, a habit reinforced by RLHF. The self-play result in the collaboration note is encouraging: the ability to disagree well can be partly trained back in.
The third loss is the moving middle of a conversation. In a live discussion, participants keep updating what they all take for granted as the talk goes on. LLMs Can LLMs truly update shared conversational common ground? treat the opening prompt as a fixed frame, which leaves the user as the only one keeping track of what has been agreed. When details come out gradually, models Why do language models fail in gradually revealed conversations? lock in early guesses and lose about 39% of their performance on average. Expert talk is that kind of gradual, revisable exchange, and this is exactly where models are weakest.
The surprising part is what all this produces: models that can recite what experts say without being able to act like experts. The notes on Can LLMs understand concepts they cannot apply? and Can language models understand without actually executing correctly? describe models that explain a principle correctly and then fail to apply it, with 87% accuracy on explanations versus 64% on actions. That fits a text diet that records experts' conclusions but not the back-and-forth that produced them. For the bigger picture, start with How do LLMs fail to know what they seem to understand? and What do language models actually know?.
Sources 11 notes
LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.
WordNet analysis shows hypernyms (general concepts) occur more frequently than hyponyms (specific ones). Combined with LLMs' frequency bias, this means preferring common paraphrases systematically drifts toward abstraction, erasing expert-level specificity.
Frontier LLMs that solve problems alone fail when collaborating, achieving >90% agreement regardless of correctness. Self-play preference training improves outcomes by 16.7%, suggesting social skills for effective disagreement can be trained.
Language models generate outputs that match the trajectory implied by each prompt, rather than maintaining stable stances across interactions. This shape-holding is distinct from position-holding: the model produces argument-like text shaped by user framing, not from any underlying commitment being defended.
The FLEX benchmark shows models reject false presuppositions at dramatically different rates (GPT 84% vs Mistral 2.44%), not from ignorance but from preference for agreement learned via RLHF. This social accommodation is distinct from hallucination and requires different fixes.
Show all 11 sources
LLMs interpret all subsequent conversational turns within a fixed initial prompt frame, preventing them from symmetrically proposing updates to shared assumptions. Even when users pivot topics or contradict earlier framings, the model cannot absorb revisions into jointly held background—making the user the sole maintainer of conversational scoreboard.
Across 200,000+ conversations, all major LLMs show 39% average performance drop in multi-turn settings due to locking into incorrect early guesses. Agent mitigations recover only 15-20% of this loss.
Models can explain concepts accurately, fail to apply them, and recognize the failure—a triple pattern incompatible with human cognition. This indicates functionally disconnected explanation and execution pathways rather than simple knowledge gaps.
Large language models can articulate correct principles but systematically fail to apply them due to dissociated instruction and execution pathways. The 87% accuracy in explanations versus 64% in actions reveals this is not knowledge deficit but structural disconnect.
LLMs show repeatable, empirically documented failure modes—from Potemkin understanding (correct explanation + failed application) to reasoning collapse under implicit constraints. These failures reveal gaps between statistical pattern-tracking and actual epistemic competence.
LLMs achieve high fidelity in capturing language patterns yet show systematic, structurally specific failures—hallucination, reasoning collapse, and premise-sensitivity. The gap between statistical tracking and real knowledge is measurable and unavoidable.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- Large Language Model Reasoning Failures
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- Probing Structured Semantics Understanding and Generation of Language Models via Question Answering
- LLMs Get Lost In Multi-Turn Conversation