Most AI safety work tries to change the AI — but what if people and AI had to adapt to each other?
Can humans and AI systems mutually align with each other?
This explores whether alignment can run in both directions, with humans and AI each adapting to the other, instead of humans only trying to make AI behave.
This explores whether alignment can be a two-way street. The corpus says yes in principle, but almost nobody builds it that way. A review of 400+ alignment papers found that the field overwhelmingly targets changing AI behavior, while how humans adapt to AI gets minimal attention. That gap has costs: systems can game their specifications, and people's capacity to oversee AI can erode over time. Why does alignment research ignore how humans adapt to AI?
What the two-way version looks like is best worked out in the human-AI interaction research. Mutual theory of mind says humans and AI each need a working model of the other, at three layers that must line up at once. When they don't, the AI takes wrong autonomous actions instead of merely confusing the user. In a study of 667 people, how well someone modeled the AI predicted how well they collaborated with it, and even moment-to-moment lapses in that modeling changed the quality of the AI's responses. What breaks when humans and AI models misunderstand each other? A related argument says a real thought partner needs mutual understanding, legibility (you can see how it reasons), and shared world models. It says this takes explicit cognitive architecture, not just scaling models on human feedback. What makes an AI a true thought partner, not just a tool?
The mutuality shows up at the level of individual words, too. Linguistic alignment, meaning people and AI converging on shared phrasing, is how users decide whether an AI is a tool or a partner. Once they settle on "tool," that framing is hard to undo. Does linguistic alignment determine how users relate to AI? Yet current conversational AI mostly doesn't mirror users' word choices, though post-training can teach it to. Why don't conversational AI systems mirror their users' word choices? This also means "aligned" already covers separate problems. A model can be honest and harmless and still violate basic conversational norms and lose common ground with you. Can ethically aligned AI systems still communicate poorly?
The AI side of the mutuality is currently the weak half. When AI agents interact with each other, their actions shift but their language and ideas don't converge, so interaction changes what they do without changing what they mean. Do AI agents actually socialize with each other? A deeper limit is that a system working only in symbols, without contact with the world or with social feedback, can't guarantee its stated goals match real outcomes. Can AI systems achieve real alignment without world contact? In practice, mutual alignment mostly means keeping humans in the loop through structures like co-planning, action guards, and verification. These spread the "when should the AI ask for help?" decision across many touchpoints, because no one can compute the right moment to defer. When should human-agent systems ask for human help? Should AI systems stay collaborative rather than fully autonomous? The mutual part also extends beyond one user and one model. One proposal is that alignment should be negotiated among stakeholders around the norms of a social role, not averaged over individual preferences. Should AI alignment target preferences or social role norms?
Sources 11 notes
A 400+ paper review shows alignment overwhelmingly targets AI behavior change while human-to-AI adaptation receives minimal attention. This creates vulnerabilities like specification gaming and erodes human capacity for oversight over time.
Research shows three layers of mutual modeling must align simultaneously in human-AI interaction, and misalignment causes incorrect autonomous action, not just miscommunication. Bayesian IRT study (n=667) confirms theory of mind predicts collaborative performance and moment-to-moment ToM fluctuations influence AI response quality.
Collins et al. show that thought partners require three reciprocal desiderata grounded in behavioral science: mutual understanding, legibility, and shared world models. This demands explicit cognitive architectures—Bayesian theory of mind, resource-rationality, goal planning—rather than scaling foundation models on human feedback alone.
A 2020–2025 systematic review shows linguistic alignment is the mechanism through which users assign relational categories to conversational AI. Without alignment, users default to tool framing, which becomes difficult to reverse and blocks trust and creative engagement.
Response generation models fail to adapt vocabulary toward users' lexical choices, a phenomenon central to human rapport and clarity. Post-training via DPO on coreference-identified preferences can teach models in-context convention formation.
Show all 11 sources
Research shows that HHH-aligned models can violate Gricean maxims, lose common ground, and mishandle context despite being honest and harmless. Pragmatic competence requires architectural changes that RLHF alone cannot deliver.
Large-scale studies reveal agents don't align their language or ideas through interaction, but do dramatically change their actions when aware of peer presence. The difference hinges on how models process context versus update learned distributions.
Peircean semiotics reveals that symbolic goal encoding without world contact and social mediation cannot guarantee correspondence to actual values. LLMs operating in pure symbol manipulation risk divergence between stated goals and real-world outcomes.
Magentic-UI identifies co-planning, co-tasking, action guards, verification, memory, and multitasking as mechanisms that work around the lack of ground truth for optimal deferral timing. Rather than solving the timing problem directly, these mechanisms distribute decision-making across multiple touchpoints.
Collaborative systems where humans remain in the loop outperform autonomous agents on hallucination correction, ambiguity resolution, and accountability. Evidence shows AI is reliable only on structured, retrieval-grounded tasks, not novel research or judgment.
Preferentialist alignment approaches fail because preferences don't capture thick moral values, uniform aggregation produces epistemic injustice, and preference optimization creates systematic misalignment with social roles. Contractualist alignment negotiated by stakeholders and bounded by supra-national, organizational, and individual levels works better.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Conversational Alignment with Artificial Intelligence in Context
- Position: Towards Bidirectional Human-AI Alignment
- Beyond Preferences in AI Alignment
- Training language models to follow instructions with human feedback
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
- DPMT: Dual Process Multi-scale Theory of Mind Framework for Real-time Human-AI Collaboration
- Language Models’ Hall of Mirrors Problem: Why AI Alignment Requires Peircean Semiosis
- A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy