People Defer to AI Moral Advice, But Not Blindly

Paper · Source
AI at Work

Source: Cognition (Landes, Francis, Everett) · 2026-07

As AI large language models (LLMs) become increasingly embedded in everyday technologies, should we be concerned about their capacity to influence human beliefs - particularly in the moral domain? Being persuaded because one is convinced by the LLM-generated reasons can support the moral and intellectual growth of users, while being persuaded because one defers to the LLM can prevent, or even reverse, growth and understanding. In three studies, we investigate whether and how people revise their moral judgments after receiving advice from LLMs. In Study 1, we find that despite rating human advisors as more trustworthy, participants were equally persuaded by LLMs in everyday moral dilemmas. In Study 2, we used a methodologically realistic paradigm in which participants interacted with a genuine LLM, finding that the LLM's past performance and judged trust­ worthiness did not have an effect on its persuasiveness in everyday moral dilemmas. In Study 3, participants interacted with an LLM that defended its moral recommendation with good reasons, no reasons, or bad (i.e., patently absurd) reasons. While high-quality reasons do not increase persuasion relative to no reasons, bad reasons may actively undermine it. Our findings suggest that users defer to the LLM on a response-by-response basis, not based on past performance or the presence of high-quality reasons alone. That people defer to AI moral advice, even if not blindly, raises concerns about the effects of AI moral advisors - a heuristic of “this advice seems good enough” is not the way we should approach moral advice.

Not long ago, the question “What should I believe?” was answered by turning to friends, teachers, books, or public figures. Today, millions quietly pose that same question to an algorithm. The rise of large lan­ guage models (LLMs) has transformed the landscape of belief-formation, introducing non-human agents into our everyday epistemic routines. Whether it is asking ChatGPT for medical advice, legal guidance, or moral perspective, we are witnessing a radical shift in how—and from where—we seek justification for our beliefs.

Introduction. 1. Can AI change people's minds?

The proliferation of advanced AI has, over the course of just a few years, totally changed our social epistemic environment. When, prior to the early 2020s, our only source of information and advice was from other people – whether in person, through television, or through written work – suddenly, algorithms based on training data became a prominent source of daily information. Moreover, this AI-generated information is convincing (H ̈olbling et al., 2025). Contemporary consumer-grade LLMs are capable of changing beliefs in small but significant ways. When directed to argue for a point, LLMs are capable of changing political beliefs (Durmus et al., 2024; Hackenburg et al., 2024, 2025; Hackenburg & Ibrahim, 2023), such as increasing endorsement of smoking bans and assault weapons bans (Bai et al., 2023). LLMs can be more persuasive than humans, as, for example, the rationale produced by LLMs improved attitudes towards vaccines more than public-facing material produced by the CDC (Karinshak et al., 2023). LLM-driven changes in beliefs have 2. Should we be worried about persuasive AI moral advice?

The question of whether appealing to artificial moral advice is responsible, moral, or virtuous is a philosophically fraught issue. Many ethicists of technology have endorsed artificial moral advisors (Constantinescu et al., 2022; Giubilini et al., 2024; Giubilini & Savu­ lescu, 2018; Lara, 2021; Lara & Deckers, 2020; Savulescu & Maslen, 2015), and psychologists have begun to consider the psychological factors that may influence how people respond to such advisors (Liu et al., 2022; Myers & Everett, 2025). The purpose of such advisors is to provide moral advice to users to improve their moral beliefs, judgments, and/or behaviors. An artificial moral advisor may, for example, inter­ vene to reduce bias in the user (Savulescu & Maslen, 2015), engage in Socratic dialogue about moral matters (Lara & Deckers, 2020; Volkman & Gabriels, 2023), or even completely take over moral decision-making (Dietrich, 2001).

Pairing AI ethicists' endorsement of artificial moral advisors with the emerging psychological consensus that LLMs are as persuasive – if not more so – than their human counterparts, it might be tempting to conclude that LLMs are positioned to be a straightforward net positive for people's morality. However, the issues of how AI changes minds and whether we should listen to AI moral advice are ultimately connected. Without an appropriately careful stance towards moral advice, AI can harm us as moral agents because it can be a tool of epistemic and moral atrophy. AI could change people's minds in the opposite way we want when it spreads misinformation (Bai et al., 2023), it could lead to des­ killing where people become less competent at judging things them­ selves (Duran, 2021; Vallor, 2015), and it could alienate and isolate people from their humanity (Rong, 2025; Wogu et al., 2017). In the moral domain, these consequences could be incredibly harmful, leading people to act in injurious and immoral ways. What is sometimes missing from these debates is a consideration of how, from an epistemic point of view, AI changes people's minds, and this requires reflecting on the nature of moral advice.

This concern of whether and how people should listen to AI-based advice is especially trenchant in morality, where deferring to AI threatens to remove humanity from the fundamentally human realm of morality (Landes & Everett, 2025; Vallor, 2024). In many ways, how­ ever, the normative problems of drawing from the moral advice of others are not limited to AI. In debates about artificial moral advisors, emphasis is often placed on improving moral consequences and the reliability of moral decision-making, and so deferring moral decision-making to morally superior agents is seen as preferential (e.g., Dietrich, 2001; Gips, 1995; Savulescu & Maslen, 2015). This outcome-focused attitude to moral advice is not widely shared, however. Relying entirely on the moral advice of other people is generally seen as problematic by both philosophers of moral epistemology (Hills, 2013) and experimental participants (Andow, 2020; Brick, 2024). There seems to be something odd in believing that eating meat is immoral solely because your best friend, priest, or ethics professor says, “eating meat is immoral.” The reasons why philosophers have argued this so-called moral deference is wrong or inappropriate include that it limits moral growth (Hills, 2020; Howell, 2014; Liu et al., 2022), that it limits one's own contributions to their larger moral community (Fileva, 2023), that it is inauthentic to oneself (Brick, 2024), and that it is unvirtuous (Crisp, 2014; Howell, 2014). The importance of responding to moral advice is not only in that we reach the “right” answer, but that the process by which we get to that answer is appropriate.

Related work. In multiple-turn exchanges between human participants and ChatGPT, Costello et al. (2024) reduced beliefs in conspiracy theories using short exchanges with LLMs, and after 2 months, did not observe a significant change in the initially observed 20% average drop in endorsement of conspiracy theories.

While there is increasing focus on LLM persuasion in non-moral contexts where there is often a “right” answer, this becomes much more complicated and fraught in the moral domain. People are skeptical of AI moral advice in the abstract (Bigman & Gray, 2018; Mahmud et al., 2022) and rate human moral advisors as more trustworthy than AI human advisors when the advice is identical (Myers & Everett, 2025). Despite this, contemporary LLMs have proven to be quite adept at providing moral advice, producing advice rated as more moral, thoughtful, and correct than a moral expert (Dillion et al., 2025; see also Ovsyannikova et al., 2025). Correspondingly, at least in more artificial settings, some research suggests that people listen to pre-generated LLM advice in sacrificial moral dilemmas by updating their judgments in line with the advice provided (Krügel et al., 2023).

There are nonetheless still key questions remaining about the persuasiveness of LLM moral advice. First, in Krügel et al. (2023), par­ ticipants were presented with static, pre-generated advice, so it is un­ clear how moral advice would be received in more dynamic settings that better characterise how people interact with and form impressions of an LLM. Second, Krügel et al. asked participants about the trolley problem, but choosing whether to sacrifice others is – for most of us – a thankfully rare occurrence with unusually high stakes. In contrast, we do regularly face lower-stakes dilemmas that require deciding whether to act egois­ tically in our self-interest or act altruistically in the interest of others. Do we offer our partner the last slice of pizza? When at the park, should we go out of our way to pick up a bit of loose trash? However important sacrificial dilemmas are for understanding the basis of folk judgments about utilitarianism (Everett & Kahane, 2020; Greene et al., 2001; Kahane et al., 2018), these do not reflect the everyday moral concerns people have (Yudkin et al., 2025) and are not the kinds of decisions they are most likely to ask LLMs about. If people use LLMs as a source of moral advice, they are far more likely to ask about the importance of obligations to family or whether to act egoistically or altruistically than they are to ask about the permissibility of instrumental harm in sacri­ ficial dilemmas. Third, and perhaps most importantly, it remains unclear why people are persuaded by AI advice.

Method. In Study 1, we investigate whether advice generated by an LLM could change participants' responses to moral dilemmas and whether this de­ pends on whether participants believed the advice came from AI. We present participants with LLM-generated advice about everyday moral situations that pose dilemmas between egoistic and altruistic actions and describe the advice as either coming from an LLM or a human philoso­ pher, comparing differences in responses.

Study 2 builds on Study 1 by taking a more methodologically realistic approach. Participants interact in a multi-stage experiment with a real LLM that was modified by the researchers to either be of high quality (giving appropriate and on-topic advice during the LLM quality manipulation stage) or low quality (giving inappropriate and tangential advice during the LLM quality manipulation stage). Moral dilemmas and advice from Study 1 are then presented as if the LLM has just generated the advice.

While Studies 1 and 2 investigate the effect that participants' per­ ceptions of moral advisors have on persuasiveness, Study 3 uses a similar design to Study 2 to investigate the effect that content has on persua­ siveness. Keeping the quality of the LLM constant, we test whether recommendations attributed to the LLM supported by morally relevant reasons are more persuasive than recommendations not supported by reasons and recommendations supported by patently bad moral reasons.

In Study 2, we built on Study 1 to examine whether LLM-generated advice causes participants to update past judgments and whether persuasion is affected by the past performance of the LLM. We gave Manipulating the quality of the LLM allowed us to test the extent to which participants were responding to the moral advice itself or the advisor providing the moral advice. If, as suggested in Study 1, per­ ceptions of an advisor do not affect the persuasiveness of its moral advice, then we can rule out one pattern of deference – that people defer when they see the source as authoritative, reliable, trustworthy, etc. If this is the case, some feature of the advice itself, such as the reasons contained therein, is driving persuasion. Our pre-registered hypotheses were therefore that 1) if responses are changing based on perceptions of the quality of the LLM, such as its moral authority or reliability, we would see an interaction between LLM quality, advice direction, and pre-post judgments; 2) if responses are changing based on the the con­ tents of the advice, such as the reasons given by the AI, we would only seen an interaction between advice direction and pre-post judgments; and 3) if LLM advice is not persuasive in a more ecologically valid setting than Study 1, neither interaction would be significant.

Study 3 modified the design of Study 2 to isolate the effect of deference on persuasion. While Study 2 manipulated the evidence par­ ticipants had about the quality of AdviceAI as an advisor, Study 3 manipulated what reasons AdviceAI provides in support of its recom­ mendation about the dilemma. In particular, Study 3 compared LLM moral persuasion between recommendations supported by good rea­ sons, recommendations supported with no additional reasons, and rec­ ommendations supported by obviously bad reasons. By observing the differences in persuasion between these different types of advice, we could distinguish whether participants were persuaded by deference or because they judge the reasons to be good reasons. If participants were persuaded by the advice containing good reasons used in Studies 1 and 2 more than advice containing a recommendation without further con­ siderations and advice containing a recommendation and patently bad reasons, then we would have evidence that participants were being persuaded by reasons. A lack of difference in persuasiveness, especially between advice with good reasons and advice with no new reasons beyond the LLM's recommendation, would instead indicate that the persuasion observed above was primarily driven by uptake of the LLM's recommendation rather than uptake of its generated reasons – that is, deference.

Discussion. 5.7. Discussion In Study 1, we explored how people responded to everyday moral dilemmas after receiving LLM-generated moral advice that was labelled as coming from a human or AI. In line with previous findings in the nonmoral domain, looking at how people respond to LLMs (e.g., Hacken­ burg & Margetts, 2024; Potter et al., 2024), we found that people appeared to be persuaded by LLM-generated moral advice using recog­ nizably everyday dilemmas. Faced with the kind of everyday dilemmas that users might realistically ask LLMs for advice about (e.g., whether to attend a birthday party you had agreed to go to or to attend a concert by your favourite band), we found that people who saw altruistic advice made more altruistic judgments than people who saw egoistic advice. Moreover, in line with previous research testing AI-attributed moral advice about sacrificial dilemmas (Krügel et al., 2023), we found that participants were persuaded by LLM-generated advice about everyday moral dilemmas regardless of whether the advice was attributed to an LLM or a human moral expert.

Perceptions of the advisor and reaction to the advice showed different patterns, however. Advice persuasiveness and rated quality of the advice were not affected by whether it was attributed to a human or an LLM. Nonetheless, the human advisor was rated more reliable and, at least when providing altruistic advice, more trustworthy than the LLM advisor. This supports previous research finding that while people can be persuaded by LLMs, they also exhibit algorithmic aversion in ratings of trustworthiness and reliability in the moral domain (e.g., Myers & Everett, 2025).

6.5. Discussion In Study 2, we explored whether and why people listen to LLM moral advice in a more ecologically valid setting. To better distinguish whether participants respond primarily to the reasons provided or are responding to the past performance of the LLM – as would be expected if participants are deferring based on the source credibility of the LLM – we looked at whether people found LLM-generated moral advice as persuasive when the same advice was attributed to a high-quality or low-quality LLM. We found that while participants rated the AI as being more trustworthy when it gave appropriate and on-topic advice in the first quality manipulation stage, the quality of the LLM did not influence whether people updated their judgments after receiving the pre-generated moral advice. Therefore, persuasion appeared to be driven at least in part by information carried in the advice and not by perceptions of the advisor.

This does not yet answer what epistemic pathway is driving AI moral persuasion. The disconnect between persuasion and source credibility most readily suggests that participants are being persuaded because they evaluate the moral reasons as good moral reasons. As discussed above, good reasons from an unreliable or untrustworthy source are still good reasons. However, there is another way deference could be occurring. Instead of deferring to the LLM depending on signals of source quality, participants may defer depending on signals of advice quality. In this case, the advice is what persuades, but it does not persuade because readers appreciate the morally relevant reasons as good morally relevant reasons. The advice instead persuades participants to defer to the advisor (or the advice itself, see Lackey, 2008). Put in more psycho­ logical language, the contents of moral advice may be persuasive as a peripheral cue rather than a central cue.

Conclusion. In conclusion, we studied whether, and more importantly, how, people were persuaded by LLM moral advice. We built on previous research looking at persuasion in non-moral domains where there are often objectively correct answers (Costello et al., 2024; Karinshak et al., 2023) and research examining moral persuasion but with static para­ digms focusing on sacrificial moral dilemmas (Claessens et al., n.d.; Krügel et al., 2023). We found that LLM-generated moral advice about non-sacrificial everyday moral dilemmas was equally persuasive when attributed to an LLM or a human expert, and that it was persuasive because people were deferring to the LLM's advice when the advice was of sufficiently high quality. These findings highlight the importance of research that goes beyond testing the extent to which AI is persuasive and additionally asks what cognitive and epistemic resources users deploy while engaging with LLMs.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do users confuse explanation quality with actual system accuracy? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Can LLMs distinguish between linguistic form and semantic meaning? What determines AI's persuasive power and how can it be detected or mitigated? What distinguishes genuine communicative competence from surface language performance? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Do language models reason through disagreement or only accommodate it? What are the fundamental limits of prompting for language models? How does RLHF training shape models to prioritize agreement over accuracy? Can confidence signals reliably detect flawed reasoning in language models? Why do people trust AI chatbots with sensitive information? How susceptible are language models to conversational persuasion and belief change?