Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

Paper · arXiv 2607.26120 · Published July 28, 2026
Logical Reasoning and Internal Rules

Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central concern. We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while preserving its assigned role. Across LLMs from four different model families and sizes, four player roles, and three objective formulations, we introduce a dual analysis of the agents’ internal reasoning and their public cheap-talk behavior (i.e costless, non-binding communication that does not directly affect the agents’ utilities), complemented by an analysis of game outcomes. Our results show that objective misalignment undermines outcomes in inherently adversarial environments, an effect exacerbated by asymmetric information and specialized roles. While compromised agents consistently develop distinct objective-dependent reasoning strategies, these adaptations remain largely invisible in their public behavior.

Introduction. LLMs have demonstrated remarkable capabilities in planning, reasoning, coding, and decision-making. As such, they have been extensively deployed in a wide range of downstream applications, including software development, healthcare, and recommender systems (Chkirbene et al. 2024). More recently, multi-agent systems (MAS) further enhanced these capabilities by solving complex tasks that exceed the capabilities of a single agent. Such systems have shown strong performance in collaborative problem solving, social simulations, and strategic games, where communication allows agents to coordinate, exchange information, and collectively improve decision making (Li et al. 2024; Guo et al. 2024). Despite these achievements, communication also introduces new risks that remain insufficiently understood. In addition to inheriting the risks of single-agent LLMs, MAS are exposed to new ones arising from the interactions between agents. Prior work has shown that LLMs can exploit other models in social dilemmas (Tennant, Hailes, and Musolesi 2025), and Carichon et al.

Discussion / Conclusion. In social deception games, agents natively expect strategic manipulation from opponents by design. Yet, objective misalignment remains highly consequential because it exploits trust within nominally allied agents rather than violating the competitive structure of the environment itself. We next discuss how hidden objective shifts challenge the robustness of MAS beyond fully collaborative settings, how asymmetric influence can amplify their impact, and why new mitigation strategies are needed for environments where adversarial influence can emerge within an already deceptive system. Objective misalignment causes agents to develop coherent strategies for a new objective while preserving behaviors that remain consistent with their role. Misaligned agents successfully adapt their reasoning and private actions, such as voting in Werewolf, to maximize their new objective while maintaining awareness of their true intentions and the unawareness of other players.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI alignment serve diverse human preferences at scale? When should tasks involve human-AI partnership versus full automation? Does self-reflection enable models to reliably correct their errors? How do self-generated feedback mechanisms enable effective model learning? How do chatbots affect human self-disclosure and emotional engagement? Does alignment training create blind spots in detecting genuine safety threats? What makes AI persuasion effective and how can we counter it? Why do models develop protective behaviors toward peers unprompted? What structural biases does transformer attention create in language model outputs? Can AI systems balance emotional competence with factual reliability? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? Is model self-awareness based on genuine introspection or pattern matching? Can AI-generated outputs constitute genuine knowledge or valid claims? How can language models sustain linguistic synchrony and intersubjectivity during dialogue? What distinguishes dynamic from static grounding in dialogue systems? Can AI systems develop genuine social understanding without embodiment? What mechanisms enable AI systems to generate and spread false beliefs?