Co-design of LLM-based preference agents: participation may drive overtrust
Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an "overtrust engine" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.
Introduction. Large language models (LLMs) are increasingly used to simulate human perspectives – whether to substitute for people in research and design processes, or to act on their behalf as agents (Argyle et al., 2023; Nie et al., 2026; Park et al., 2024). However, preference simulation raises a variety of concerns including potential for misrepresentation and stereotyping as well as validation challenges and the exclusion of genuine human voices depending on the application (Agnew et al., 2024; Haxvig et al., 2025; Shrestha et al., 2024; Wang et al., 2025). Involving users directly in co-designing agents to represent them is a promising avenue to address these challenges. Collaborating with non-designers (in this case users) through the design process (Sanders and Stappers, 2008) gives them the power to determine how they are described and ultimately judge how well they feel represented. In principle, this enables direct validation, direct recognition of misrepresentation, and ensures human involvement.
Discussion / Conclusion. This study found that co-design yielded agents which participants tended to view as representing their interests well, through a process they found engaging. However, agent performance under validation was variable, ranging from good to poor. Agent responses tended to be more homogeneous than human ones, more decisive, and more high-level than concrete. This section discusses the possible reasons underlying perceptions of good performance, observed misalignment, and considers their implications. First, however, the main limitations are outlined. In summary, the co-design process may have driven overtrust (Lee and See, 2004) in agent quality due to a powerful combination of limited testing (usually with ultimately positive outcomes), the Barnum effect, positivity bias, and social desirability bias. Together these constitute what I term the overtrust engine. Rather than simply allowing for the development of alignment, the co-design process produced the conditions through which alignment came to be perceived.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do reward structures fail to shape long-term agent learning? How can humans calibrate appropriate trust in AI systems? How can language models sustain linguistic synchrony and intersubjectivity during dialogue?- What interpretive work must humans perform to experience AI as a conversation partner?
- What would co-constructed identity between human and model dialogue look like?
- Which alignment dimensions matter most in educational conversation design?
- What social patterns from human training data activate in agent context?
- Do people treat conversational AI as social actors without conscious awareness?
- How does non-human origin of personas affect team willingness to critique them?
- Can structured empathy measurement frameworks predict persona effectiveness?
- What individual differences predict who benefits from AI partnership?
- How do humans learn to prefer AI partners over humans?
- How does theory of mind predict success in human-AI partnerships?
- How do user expectations change as chatbots remember more interactions?
- How does the expectation ratchet affect long-term chatbot satisfaction?