SYNTHESIS NOTE
Topics›Argumentation›this note

Do fluent arguments win debates through sound logic or rhetorical polish?

When LLMs and humans debate, do subjective judges and formal logic reach the same verdict about who argued better? Testing both measures jointly reveals whether winning an argument depends on logical rigor or persuasive framing.

Synthesis note · 2026-09-25 · sourced from Argumentation

The paper ranks the same debaters two ways and gets two different orderings. Its PERSUASIO platform is "grounded in a formal argumentation-based theory of persuasion dialogues" and adjudicates logical winners in free-text debates. The authors generated 192 debates on a UK political topic between humans and LLMs, then evaluated 22 interlocutors through automated adjudication and 9,702 crowdsourced pairwise judgements. LLMs "dominated the subjective ranking yet performed substantially worse under argumentation-theoretic adjudication, where humans remained competitive." The discussion adds "pronounced rank inversions among top-performing subjective systems" and names the result "a systematic gap between rhetorical fluency and formal argumentative strength."

The paper's framing is that existing evaluations assess either rhetorical appeal through human judgment or argument structure through automated analysis, "but rarely both jointly," so no one has checked whether "model fluency and confident framing align with inferential rigour and logical consistency." Putting both measures on the same debates is what exposes the split. In the authors' reading, logical adjudication "rewards different properties than those that drive human perceptions of persuasiveness," while the LLMs studied appear "optimised for rhetorical polish – fluency, structural clarity, confident framing." Two causes are offered, both explicitly tentative: autoregressive generation "perhaps" leaves models myopic, unable to plan over upcoming content, and instruction tuning and RLHF are "potentially" compounding this by optimizing for fluency. Multi-agent orchestration and retrieval augmentation "seemingly amplify" the divergence, raising perceived persuasiveness "without reliably safeguarding logical integrity."

This is a second, independent route to the separation described in Can LLMs persuade without actually understanding arguments?. That note shows LLMs failing to evaluate debates they can win; here a formal adjudicator scores the LLMs' own arguments and finds them weaker than their reception suggests. "Confident framing" sits on the rhetorical side of the paper's ledger, which fits Does linguistic conviction explain why LLMs persuade more effectively?, though the excerpt does not measure conviction. The post-training conjecture parallels Where does AI's persuasive power actually come from?, with logical rigor standing in for factual accuracy. The multi-agent and retrieval result echoes When does debate actually improve reasoning accuracy?: added machinery raised the persuasive surface without adding a check on soundness.

The excerpt does not name the models, give the size of the rank shifts, or say what the adjudication theory scores beyond "formal coherence." It covers one political topic and does not describe measuring belief change, so it does not directly conflict with Are language models actually more persuasive than humans?. The autoregressive and RLHF explanations are conjecture; the excerpt reports no test of either. What the evidence does support is narrower: in this setup, how persuasive an LLM debater sounds to human raters is a poor guide to how well it argues under formal adjudication, and an evaluation that relies on human preference alone will place LLMs higher than an argumentation-theoretic one would.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do LLM judges' systematic biases affect alignment and evaluation outcomes? What factors drive AI persuasiveness and how can it be mitigated? Can multi-agent systems avoid converging on false agreement without deliberation?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 79 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLMs dominate subjective persuasiveness yet fall sharply under argumentation-theoretic adjudication — fluency is not formal argumentative strength