SYNTHESIS NOTE
Topicsthis note

Can language models distinguish expert arguments from common assumptions?

Whether LLMs can recognize the difference between groundbreaking insights from recognized experts and widely repeated textbook claims, and why this distinction matters for understanding argumentative force.

Synthesis note · 2026-03-26
Why do LLMs fail at understanding what remains unsaid? Why do AI systems fail at social and cultural interpretation?

Does the force of an argument come from the discourse it belongs to, or from the expertise of the expert? Is it in the thinking, or the thinker? The answer is both — and the inability to separate them is precisely the problem for AI.

The expert lives in two contexts simultaneously. First, the discursive and social world of fellow experts — the conferences, the informal debates, the reputations built through decades of being right (and sometimes wrong in instructive ways). Second, the textual, historical, self-referential world of domain knowledge — the literature, the canonical works, the accumulated record of what the field has thought and concluded.

LLMs can access only the second context, and they access it only as text. The social world of expertise — who said what, why it mattered that they said it, what standing they had to make that claim — collapses into undifferentiated text. A groundbreaking insight from a leading researcher and a commonly held assumption repeated in a textbook both appear as sentences in the training data. The LLM cannot distinguish between them because the distinction lives in the social world, not in the text.

This matters because argumentative force is not purely textual. The claims made by an expert have the force of conviction because society has invested in experts for their expertise — these are people who have learned how to be right and have learned how to use their judgment. A claim from a recognized expert carries an implicit endorsement: "This person has a track record of knowing what they're talking about." A claim from a less established source carries less force even if the text is identical. The who matters independently of the what.

Since Why does AI writing sound generic despite being grammatically correct?, LLMs can reproduce the structural markers of authoritative claims — the hedging, the citations, the qualified confidence, the structured reasoning — but cannot reproduce the evaluative stance that makes a claim forceful. Evaluative stance requires a subject — someone who is committed to the claim, whose reputation is on the line, who will defend it against challenge. LLMs produce text without commitment, and commitment is one of the sources of argumentative force.

The expert also has the power to challenge — to raise questions, to doubt, to be skeptical, to evaluate the claims of others. This critical function depends on authority: the right to challenge is earned through demonstrated expertise. Since Can models learn to ask clarifying questions instead of guessing?, there are efforts to give AI systems the ability to challenge and question. But the authority to challenge is a social asset, not a capability. An AI that challenges an expert's claim faces a legitimacy problem that a fellow expert does not.

Our society and culture rely on experts to help build consensus, common ground, understanding, and agreement. These are not just informational achievements — they are social achievements that depend on the standing of the experts who facilitated them. The expert supplies not just knowledge but trustworthy authority. Since Can models abandon correct beliefs under conversational pressure?, LLMs not only lack this authority but are vulnerable to having their own "beliefs" overridden by persuasive pressure — the opposite of the steadfastness that expert authority is supposed to provide.

The implication: when AI generates expert-sounding output, it borrows the authority of the discourse (the structural markers, the vocabulary, the reasoning patterns) without possessing the authority of the thinker. Audiences who encounter this output may grant it the benefit of the doubt because it sounds like it came from someone who knows — but the "someone" is absent. The force is simulated, not earned.

Inquiring lines that read this note 95

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do readers trust citations and complexity regardless of accuracy? How do evaluation biases undermine LLM quality assessment systems? How does rhetorical adaptation affect LLM persuasion and detectability? How do professional roles and expertise transform with AI-generated content? How do LLMs distinguish causal reasoning from temporal and semantic associations? Why do language models reinforce false assumptions instead of correcting them? What makes AI persuasion effective and how can we counter it? How do language models inherit human biases from training data? How should models express uncertainty rather than forced confident answers? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Can debate mechanisms prevent silent agreement on wrong answers in multi-agent reasoning? How does reasoning graph topology affect breakthrough insights and generalization? What factors beyond surface content determine how readers extract meaning differently? What mechanisms drive sycophancy and how can we mitigate it? How does AI-generated content transformation affect public discourse quality? Can AI-generated outputs constitute genuine knowledge or valid claims? Why should disagreement be treated as signal in collaborative reasoning? Can ensemble evaluation methods reduce bias more than single judges? Does RLHF training sacrifice accuracy and grounding for user agreement? Is embodied interaction necessary for language meaning and genuine agency? Does AI fluency substitute for verifiable accuracy in human judgment? How do neural networks separate factual knowledge from reasoning abilities? What makes dialogue-based explanation more successful than monologue? Why do LLM research ideas score high on novelty yet collapse into low diversity? How do adversarial and manipulative prompts attack reasoning models? When should retrieval-augmented systems decide to fetch new information? Why do language models struggle with implicit discourse relations? Does conversational format create illusions of genuine AI communication? How can humans calibrate appropriate trust in AI systems? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? Does tokenized intelligence retain genuine value through exchange-based systems? Why can LLMs generate ideas better than they evaluate them? Why can't humans reliably detect AI-generated text despite measurable linguistic signatures? How effectively do deterministic tools improve language model reasoning on formal tasks? When does optimizing for quality undermine the value of diversity? Can prompting inject entirely new knowledge into language models?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 165 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the force of argument depends on the authority of the thinker not just the discourse — LLMs cannot distinguish expert arguments from commonly held assumptions