INQUIRING LINE

The AI explanations people like best can be the ones that make them worse at catching the AI's mistakes.

What makes a rationale interface trustworthy versus merely satisfying to users?

This explores what separates a rationale that helps people judge whether an AI is right from one that just feels good to read.


This explores what separates a rationale that helps people judge whether an AI is right from one that just feels good to read. The corpus's blunt answer is that the two often pull in opposite directions. Satisfaction is a poor stand-in for earned trust, and the features people like most can be the ones that make them worse at catching mistakes.

The clearest evidence is a controlled study where people preferred planning-and-decomposition reasoning formats, yet plain chain-of-thought traces were better for spotting errors, calibrating trust and interpreting the output. The favored formats raised false alarms and unwarranted trust (Do people prefer the reasoning formats that help them verify?). Visual rationales show the same split from another side. Argument maps improved trust calibration on verbal reasoning tasks but hurt it on visual ones, and the satisfaction and helpfulness ratings flipped in each domain. So no rationale format is trustworthy on its own. What matters is how well the format fits the task (Do visual rationales help or hurt how people calibrate trust?). Even asking users how they feel is unreliable, because people report satisfaction while still confused, especially when they can't see their own knowledge gaps (Does user satisfaction actually measure cognitive understanding?).

Part of the reason is that a rationale can look like reasoning without being reasoning. Focus-group users trusted ChatGPT because it responded promptly and in a conversational way, not because they judged its accuracy (Does conversational style actually make AI more trustworthy?). Chain-of-thought prompts with logically invalid steps performed nearly as well as valid ones, which suggests models pick up the form of reasoning rather than genuine inference (Does logical validity actually drive chain-of-thought gains?). Reasoning traces also rarely explain decisions faithfully. An influence can be missing from the trace entirely, or problematic reasoning can show up in clean-sounding language (Can we actually trust reasoning model outputs?). A fluent trace is therefore easy to like and hard to verify. Liking and trusting are also separate processes. People rated AI moral arguments higher until they learned the source was AI, and then their agreement dropped (Do people prefer AI moral reasoning when they don't know the source?).

The stakes are high because checking is costly and fluent output breeds false confidence. Studies find roughly 80% of AI outputs adopted without challenge, a pattern called cognitive surrender (When do users stop checking whether AI output is actually backed?). A rationale that makes the answer feel settled encourages that surrender.

The one design the corpus finds genuinely helpful is the contrastive dual explanation, which argues both for and against the answer. Ordinary reasoning traces and post-hoc explanations raised acceptance whether or not the answer was correct. Only the both-sides format helped users tell right answers from wrong ones (Do explanations actually help users spot AI mistakes?). A trustworthy rationale, then, is one that keeps the user's doubt switched on for the answers that deserve it. That can make it less pleasant to use, which is why satisfaction ratings alone can't tell you whether you've built it.


Sources 0 notes