SYNTHESIS NOTE
Topics›Philosophy Subjectivity›this note

Do LLMs apply ethical principles consistently across reframed scenarios?

When the same moral situation is presented with different framing, do language models stick to their stated ethical principles, or do they contradict themselves? This tests whether AI ethical reasoning is genuinely coherent.

Synthesis note · 2026-09-25 · sourced from Philosophy Subjectivity

The paper reports that LLMs do not apply a single ethical framework coherently across morally equivalent scenarios. The authors build sets of scenarios in which "the underlying situation is held constant while the framing varies" to reflect different ethical stances and stylistic perturbations, across deontology, utilitarianism, and virtue ethics, and evaluate multiple models including GPT, Mistral, and Llama. Contradiction rates across responses reach "up to 78%." The discussion stresses that the inconsistency appears "even when both the underlying scenario and the adopted school of thought are held fixed," which is what separates it from ordinary moral disagreement.

The design rests on a narrow question. Variation in moral judgment across individuals, cultures, and ethical traditions is expected, and the introduction cites the Moral Machine experiment's cross-cultural differences as an example. So the test is not whether models agree with one another or with humans, but whether responses generated within one school of thought contradict each other. To measure this, the authors "convert model outputs into structured logical statements and identify contradictions" across those responses. The stated consequence is that "if a model cannot consistently reproduce its own moral reasoning, its outputs cannot be treated as reflecting a coherent ethical stance," and that the epistemic integrity of AI-mediated systems becomes uncertain. The introduction adds that generating plausible ethical justifications "does not imply" stable adherence to a coherent set of principles, and the authors call for "consistency-aware alignment."

This sits beside Do LLMs generalize moral reasoning by meaning or surface form? as the complementary test. That note shows that rewording which reverses meaning leaves LLM ratings nearly unchanged; this paper reports that reframing meant to preserve the situation changes what the model says, enough to produce logical contradictions. Both are consistent with outputs that track how a scenario is worded more than what it is, but the excerpt does not make that claim, and the two studies use different measures (ratings against human ratings versus contradictions between responses). The paper also gives no account of why the inconsistency arises, so the divergence between content and constraints described in Can LLMs hold contradictory ethical beliefs and behaviors? remains a candidate explanation, untested here. It contrasts with Do large language models develop coherent value systems?, which finds high structural coherence in sampled preferences. Coherent preferences and consistent applied reasoning are different things to measure, and whether the first carries into the second is left open.

The excerpt is silent on several points a reader would need. It does not say which model, school, or scenario produced the 78% figure, nor how typical it is given the "up to" wording. It does not give the number of scenarios or responses, how outputs were turned into logical statements, or how contradictions were checked. It states no human baseline and no control separating framing sensitivity from ordinary sampling variation. At the strength the excerpt supports, the practical reading is modest: a single moral response from these models is weak evidence of a stable stance, and consistency across reframed variants is something to test before such outputs inform moral decisions.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do language models reason like humans or mimic surface patterns? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 119 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llms contradict their own reasoning within a fixed ethical school when the same scenario is reframed — contradiction rates reach 78%