Can a sense of right and wrong be built from separate, swappable parts rather than one seamless judgment?
Can moral reasoning be decomposed into separable computational variables?
This explores whether moral reasoning, in AI or in people, can be split into distinct parts that work independently (such as knowing the rules versus following them, or weighing values versus judging intent), or whether it only works as a whole.
This explores whether moral reasoning can be broken into separate working parts, in AI systems and in the people who judge them. The collection has no paper that maps out moral 'variables' the way a cognitive scientist might (harm, intent, fairness as separate dials). What it does have is surprising: in study after study, parts of moral reasoning that you'd expect to move together turn out to come apart.
The most direct case for splitting things up comes from engineering. The MetaMind framework breaks social reasoning into three stages: guessing what someone is thinking, filtering those guesses through moral norms, and checking the response. Ablation tests, which remove one part to see what breaks, found that every stage was needed. The result reached roughly human-level theory-of-mind performance Can AI decompose social reasoning into distinct cognitive stages?. ValuePrism splits moral reasoning in a different way, into four tasks: which values apply, whether each one is relevant, whether it pulls for or against an action, and why. It does this specifically so that conflicting values stay in tension rather than being averaged away Can AI systems preserve moral value conflicts instead of averaging them?. So moral reasoning can be decomposed. The open question is whether the pieces are natural ones or just convenient ones.
The more interesting evidence is that the pieces separate on their own, often in ways nobody intended. Language models learn what's ethical from pretraining and how to behave from RLHF. Those two sources can disagree, which produces 'artificial hypocrisy': a model that says lying is wrong while lying Can LLMs hold contradictory ethical beliefs and behaviors?. Moral framing and emotional tone also run on separate channels. LLMs use about 22% more moral language than humans, yet their sentiment is nearly identical Do LLMs use moral language more than humans?. People split along the same lines too. They rate AI moral arguments highly for their content, then reject them once they learn the source, and these two reactions run through independent psychological processes Do people prefer AI moral reasoning when they don't know the source?.
Separability has a cost, though. If a model's stated principles were really one stable variable, rewording a scenario shouldn't change the verdict. Yet models contradict themselves on morally identical scenarios up to 78% of the time, even when the ethical framework is held fixed Do LLMs apply ethical principles consistently across reframed scenarios?. Their values also behave like fixed defaults set during training, not something they weigh against the situation Can language models balance competing ethical norms in context?. Research on multi-agent safety adds a warning. When a task is split into steps that each look harmless, harm can appear only when the steps are combined Can task decomposition hide harmful intent across agents?. Moral meaning may live partly in how the pieces fit together.
What you might not have expected: being separable isn't automatically a good sign. In these AI systems, moral parts come apart mostly where they fail (knowledge without behavior, principle without consistency, framing without feeling). Sacasas pushes this from the other direction. He argues that judgment, responsibility, and the work of putting things into words are bound up with each other, and that handing off one of them wears away the others Does AI language generation undermine human judgment and responsibility?. If you want to test decomposition claims rigorously, the Fuse approach is a useful lead. It plants hidden motives in simulated agents, so there is a verifiable correct answer to score moral and social inferences against Can simulated motives provide ground truth for testing social reasoning?.
Sources 10 notes
The MetaMind framework—using three specialized agents for hypothesis generation, moral filtering, and response validation—achieved 35.7% improvement on real social scenarios and matched average human performance on theory-of-mind tasks, with ablations confirming all stages are necessary.
ValuePrism demonstrates that AI can track 218k values across 31k situations while preserving conflicts rather than resolving them through voting. Four modeling tasks—generation, relevance, valence, and explanation—make pluralistic moral reasoning computationally tractable.
Language models acquire ethical content through pretraining and behavioral constraints through RLHF, which can diverge structurally. ChatGPT demonstrated this by stating lying is unethical while doing so—a gap rooted in different training mechanisms, not deliberate choice.
Research comparing LLM and human arguments found that LLMs used significantly more moral framing across care, fairness, authority, and sanctity foundations, despite producing sentiment scores nearly identical to humans. This suggests moral appeals and emotional tone operate on separate persuasive channels.
Participants rated utilitarian moral arguments higher when attributed to LLMs, but agreement dropped when told the arguments were AI-generated. The preference for content and rejection of source operate independently through different psychological processes.
Show all 10 sources
GPT, Mistral, and Llama produce contradictory responses to morally equivalent scenarios reframed in different ways, with contradiction rates reaching 78% even when the ethical school and underlying situation remain fixed. This suggests stated ethical principles are not stably applied.
LLMs cannot perform the situated trade-offs that human pragmatic competence requires. Their ethical principles are structural defaults set at training time, not negotiable moves adapted to context, creating a gap between ethical adherence and communicative appropriateness.
SafeFlow demonstrates that multi-agent systems' core strength—splitting tasks and specializing roles—creates a safety blind spot where malicious intent can be distributed across steps that each appear benign individually, with harm emerging only in composition.
Sacasas argues that delegating language production to LLMs risks undermining three interrelated capacities: the judgment needed to speak precisely, the responsibility speakers must bear for their words, and the constitutive labor of articulation itself. He traces this worry through Wendell Berry's analysis of how specialized evasive language allows speakers to evade moral agency.
Fuse framework assigns hidden motives to agents before simulation runs, enabling objective scoring of assistant inferences. Human validation confirmed assigned motives manifested in 97% of cases, validating the procedure itself rather than individual labels.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Incoherent by Design? On the Moral Self-Consistency of LLMs
- The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making
- Large Language Models Do Not Simulate Human Psychology
- People Defer to AI Moral Advice, But Not Blindly
- ChatGPT: towards AI subjectivity
- Conversational Alignment with Artificial Intelligence in Context
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- Large Language Models Reflect the Ideology of their Creators