Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation. By evaluating failure modes across different misalignment concepts, such as deception, emergent misalignment, and sycophancy, we show how conceptual ambiguity, non-robust datasets, experimental design, and insufficient causal interventions can lead to overinterpretation of model behaviors. This position paper aims to offer guidance on evidentiary considerations that can help improve methodological rigor in AMR. To achieve this, we provide a clear call to action through a proposed framework of evidence levels and a diagnostic checklist. These shared standards will enable more productive scientific discourse and ensure that claims about AI risks rest on solid empirical foundations.
Introduction. Can I trust my AI assistant? This question becomes increasingly relevant with the rapid adoption of large language models (LLMs) and artificial intelligence (AI) agents. Recently, AI systems have advanced significantly in terms of capability and general “intelligence”, which is also reflected in the nature of their failure modes. Many frontier models display eerie “human-like” failure modes, including behaviors that resemble deceptionÛ (Park et al., 2024), schemingÛ (Meinke et al., 2024), instrumental goals (Ward et al., 2024), and more (Sharma et al., 2024; Laine et al., 2024; Schlatter et al., 2026). We refer to such failures as instances of anthropomorphic misalignment. Deploying advanced AI that exhibits anthropomorphic misalignment in high-stakes environments could have catas- trophic consequences, such as power-seekingÛ (Carlsmith, 2022) or loss of controlÛ (Bostrom, 2014; Turchin & Denkenberger, 2020).
Discussion / Conclusion. The study of anthropomorphic misalignment remains a vital pillar of AI safety, offering insights into how complex models might behave in high-stakes environments. The challenges and recommendations discussed aim not to diminish this research but to strengthen its scientific foundation. By shifting the field’s focus towards more precise target framing, diverse data construction, robust experimental design, and rigorous causal-mechanistic attribution, observations of model behavior can be grounded in reproducible and technically sound evidence. As the community moves from exploratory, pre-paradigmatic behavioral studies toward a more mature, solid science of alignment, these standards will help ensure evaluations provide the technical clarity necessary to effectively inform researchers, developers, and policymakers regarding decisions around serious AI risks.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What mechanisms enable AI systems to generate and spread false beliefs? How does latent reasoning compare to verbalized chain-of-thought? Is embodied interaction necessary for language meaning and genuine agency? How do multi-agent systems achieve genuine cooperation and reasoning? When should tasks involve human-AI partnership versus full automation?- How do humans and AI develop accurate models of each other?
- How does theory of mind predict success in human-AI partnerships?
- How does theory of mind predict who benefits from AI collaboration?
- What happens when bidirectional theory of mind between humans and AI breaks down?
- What prevents humans from adapting their behavior when competing against AI?
- Can AI systems recognize intelligence in humans the way humans recognize it in each other?
- What role does bidirectional model updating play in human-AI understanding?
- How does AI sycophancy affect users' ability to repair conflict?
- Can bidirectional model updating between humans and AI reduce misalignment?
- Which application domains like healthcare and education lack alignment research?
- What happens to human expectations when they mistake consistent AI behavior for human behavior?
- Do culturally distinct human groups create similar attribution errors as human-AI mixtures?