SYNTHESIS NOTE
Topics›Alignment›this note

Does co-design participation hide misalignment in preference agents?

When people help design AI agents to represent their preferences, do they feel the agents represent them well even when independent testing shows they don't? This matters because participation is often assumed to fix representation problems.

Synthesis note · 2026-09-25 · sourced from Alignment

A primarily qualitative study of 12 participants who co-designed personal preference agents in the domain of household energy found a split between how the agents felt and how they performed. Participants "engaged readily and mostly came to see their agents as representing them well." Independent validation, however, "revealed mixed human-agent alignment," with agent responses "markedly more homogeneous, decisive, and abstract than the human sample," and performance ranging from good to poor. The author's argument is that participation, the usual remedy for misrepresentation and exclusion in LLM preference simulation, may "mask the problems it appears to solve."

The mechanism the paper proposes is an "overtrust engine" (building on Lee and See, 2004, on trust in automation). Participation and process transparency promote trust while concealing systematic misalignment. The discussion names four contributing conditions: limited testing that usually ends in positive outcomes, the Barnum effect, positivity bias, and social desirability bias. These are offered as possible reasons for the perceived good performance, not as separately isolated causes. The framing move is to treat individual alignment "not as a fixed state but as an enacted process": the co-design process "produced the conditions through which alignment came to be perceived" rather than simply allowing alignment to develop. One reading, mine and not the paper's, is that the direction of the error matters here: a homogeneous, decisive, abstract answer is generic enough to read as recognizable to its owner, which makes it hard to catch through the light testing a co-design session allows.

This sits close to Are RLHF annotations actually measuring genuine human preferences?, which argues that elicited responses can be artifacts of the elicitation. The present paper moves the same worry from the training data to the validation step: a participant's judgment that "this represents me" is itself an elicited response, exposed to the positivity and desirability biases the paper lists. It also contrasts with Does chatbot personalization build trust or expose privacy risks?, where rising trust arrives alongside rising privacy concern. Here the excerpt describes trust rising with no visible counterweight, and the cost is hidden. Against Can AI agents learn people better from interviews than surveys?, which reports high fidelity for interview-seeded agents in a 1,052-person study, the small energy-domain study adds a caution about scope and about what counts as validation. It does not overturn that result: the populations, tasks, and validation designs differ.

The excerpt leaves a lot open. It does not report how the independent validation was scored, how many of the 12 agents fell at the poor end, or whether any of the four bias conditions was tested individually. It describes no comparison with agents built without participation, so it cannot say that co-design produced more overtrust than other methods would. The claim that misalignment carries "potential structural consequences at scale" is an argument, not a measured result. What the evidence supports at this strength is a narrower design rule: in participatory preference-agent work, participants' own sense of being represented should not be the only validation, and an independent check against a human sample should sit alongside it.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does polished presentation create unearned authority in AI outputs? How do training data properties determine the emergence of internal misalignment?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 110 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

co-designing a preference agent may drive overtrust — participants mostly felt represented while validation found agents more homogeneous and abstract