Does LLM moral advice persuade through reasoning or just recommendations?
Do people adopt LLM moral advice because they trust the source, evaluate its reasoning, or simply defer to the recommendation itself? Understanding what drives moral deference matters for how we should use AI advisors.
In three pre-registered studies on everyday, non-sacrificial moral dilemmas — the kind people are actually likely to ask an LLM about, like whether to attend a birthday party or pick up litter in the park, rather than trolley-style sacrificial scenarios — Landes, Francis, and Everett find that LLM-generated moral advice is "equally persuasive" whether participants are told it comes from an AI or from a human philosopher (Study 1), even though participants still "rate human advisors as more trustworthy." In Study 2, manipulating a real LLM's track record (giving it a history of appropriate versus tangential advice) changed how trustworthy participants rated it, but the manipulation "did not have an effect on its persuasiveness." In Study 3, holding the LLM's quality constant and varying only the reasons behind its recommendation, "high-quality reasons do not increase persuasion relative to no reasons" — a bare recommendation was just as persuasive as one backed by good reasons — but "bad reasons may actively undermine it."
The authors read this pattern as deference rather than reasoned uptake. People are not updating their moral judgments because they judge the LLM trustworthy, nor because they evaluate its stated reasons as good reasons — if either were true, quality or reason manipulations would have moved persuasion, and they didn't. Instead, "users defer to the LLM on a response-by-response basis," a pattern driven by "uptake of the LLM's recommendation rather than uptake of its generated reasons." In the paper's own borrowed psychological language, the recommendation functions as "a peripheral cue rather than a central cue": removing reasons entirely doesn't weaken it, but a reason bad enough to be "patently absurd" does. Their stated conclusion is blunt: a heuristic of "this advice seems good enough" is not the way people should approach moral advice, because deference that isn't blind is still deference, with the same risks of moral deskilling and stunted growth that philosophers have long attributed to deferring to any advisor, human or AI.
This complicates Do people prefer AI moral reasoning when they don't know the source?, which found a trolley-problem anti-AI bias: participants who preferred LLM-generated content nonetheless agreed with it less once told the source was AI. Landes et al.'s Study 1 finds no such attribution penalty for everyday dilemmas — AI and human-philosopher labels produced equal persuasion despite the same trust gap the Moral Turing Test paper documents. The two results aren't necessarily contradictory; the dilemma types differ (sacrificial trolley scenarios inviting explicit utilitarian justification versus everyday egoism/altruism tradeoffs), which suggests attribution penalties on moral persuasion are dilemma-dependent rather than a stable anti-AI bias that generalizes across all moral content. The finding also runs alongside Can sycophantic AI advice still push people away from polarized views?: both locate persuasion in something about the advice itself rather than simple deference to source reputation, but this paper's Study 3 goes further, showing that the advice's content need not be reasoned through at all to do the persuading — the bare recommendation carries the effect that reasons were assumed to carry.
The excerpt doesn't establish whether this response-by-response deference holds up over repeated interactions or at real personal stakes rather than in single-shot vignettes — the paper's own cited related work on conspiracy-belief reduction found an effect that persisted at two months, but these three studies measure a single exposure. It also doesn't explain why bad reasons undermine persuasion while good reasons add nothing beyond the bare recommendation — whether that is a floor effect, triggered suspicion, or something else is left open. If the pattern generalizes, it implies that safeguards aimed at source transparency or published AI track records may protect moral judgment less than assumed, since in this design neither source credibility nor reasoning quality is what the persuasion is actually riding on.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do users confuse explanation quality with actual system accuracy?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do people prefer AI moral reasoning when they don't know the source?
Explores whether humans genuinely prefer AI-generated moral justifications or whether source knowledge changes their evaluation. This matters for understanding whether AI reasoning quality is underestimated in real-world deployment.
contrasts: that paper's trolley dilemmas show an anti-AI attribution penalty this paper's everyday dilemmas don't
-
Can sycophantic AI advice still push people away from polarized views?
Does an AI system that flatters users and agrees with their initial leanings still manage to depolarize their choices? This matters because it challenges assumptions about how AI bias affects human decision-making.
both locate persuasion in advice content over source reputation, but this paper shows content need not be reasoned through
-
Do language models judge persuasion the way humans do?
Do LLMs recognize which arguments actually change human minds, and if not, what cues do they rely on instead? Understanding this matters for using AI in social simulations and persuasion research.
complicates the credibility cue this paper finds doesn't drive moral persuasion
-
Are language models actually more persuasive than humans?
Does the research evidence support claims that LLMs persuade more effectively than humans, or have we been cherry-picking studies to fit a narrative?
Evidence for A: meta-analysis of 7 studies finds no detectable LLM-vs-human persuasion gap, supporting equal persuasiveness
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- People Defer to AI Moral Advice, But Not Blindly
- Exploring the Role of Prior Beliefs for Argument Persuasion
- Enhancing AI-Assisted Group Decision Making through LLM-Powered Devil's Advocate
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- Incoherent by Design? On the Moral Self-Consistency of LLMs
- The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making
- PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour
Original note title
people defer to llm moral advice on a response-by-response basis, not based on past performance or reason quality — but bad reasons still undermine it