Co-design of LLM-based preference agents: participation may drive overtrust

Paper · arXiv 2607.21757 · Published July 23, 2026
LLM Alignment

Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an "overtrust engine" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.

Introduction. Large language models (LLMs) are increasingly used to simulate human perspectives – whether to substitute for people in research and design processes, or to act on their behalf as agents (Argyle et al., 2023; Nie et al., 2026; Park et al., 2024). However, preference simulation raises a variety of concerns including potential for misrepresentation and stereotyping as well as validation challenges and the exclusion of genuine human voices depending on the application (Agnew et al., 2024; Haxvig et al., 2025; Shrestha et al., 2024; Wang et al., 2025). Involving users directly in co-designing agents to represent them is a promising avenue to address these challenges. Collaborating with non-designers (in this case users) through the design process (Sanders and Stappers, 2008) gives them the power to determine how they are described and ultimately judge how well they feel represented. In principle, this enables direct validation, direct recognition of misrepresentation, and ensures human involvement.

Discussion / Conclusion. This study found that co-design yielded agents which participants tended to view as representing their interests well, through a process they found engaging. However, agent performance under validation was variable, ranging from good to poor. Agent responses tended to be more homogeneous than human ones, more decisive, and more high-level than concrete. This section discusses the possible reasons underlying perceptions of good performance, observed misalignment, and considers their implications. First, however, the main limitations are outlined. In summary, the co-design process may have driven overtrust (Lee and See, 2004) in agent quality due to a powerful combination of limited testing (usually with ultimately positive outcomes), the Barnum effect, positivity bias, and social desirability bias. Together these constitute what I term the overtrust engine. Rather than simply allowing for the development of alignment, the co-design process produced the conditions through which alignment came to be perceived.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do reward structures fail to shape long-term agent learning? How can humans calibrate appropriate trust in AI systems? How can language models sustain linguistic synchrony and intersubjectivity during dialogue? Can AI systems develop genuine social understanding without embodiment? How do interface design choices shape consciousness attribution? Does conversational format create illusions of genuine AI communication? Why do persona-level simulations fail to predict individual preferences accurately? What coordination failures limit multi-agent LLM systems as they scale? How do social dynamics and selection effects compound in rating aggregates? Why do agents confidently report success despite actually failing tasks? When should tasks involve human-AI partnership versus full automation? How do chatbots affect human self-disclosure and emotional engagement? How do language models establish social grounding in human dialogue? How should personalization be implemented to improve AI assistant effectiveness? How should conversational agents balance goal-driven initiative with user control? Why do language models reinforce false assumptions instead of correcting them?