Does an AI track 'who said this, me or you' as one clear internal signal — and can flipping it change whether it accepts being corrected?
Can a conversational role attribution direction in representation space causally gate correction?
This explores whether a model's internal sense of who said what (user versus assistant) might be stored as a single measurable direction in its activations, and whether nudging that direction would decide whether the model accepts or makes corrections in a conversation.
This explores whether a model's internal record of who said what, user or assistant, could be a single direction in its activations, and whether shifting that direction would change whether the model accepts a correction. The collection has no paper that tests this directly. Nothing here isolates a role-attribution direction or links it to correction behavior, so read what follows as the nearby evidence and a guide to how the question could be tested, not as an answer.
The most useful starting point is about method. Finding a direction that correlates with conversational role would not show that it controls anything. Can LLM understanding rely on just representation or causation alone? argues that you first locate a candidate in the representations and then intervene on it to prove it has an effect. Doing only one of those gives you either a correlation or an unexplained effect. So the question really has two parts: does a role direction exist, and does changing it change what the model does with a correction? The second part is the hard one.
There is indirect reason to think intervening on the representation could work where prompting fails. Why do language models ignore information in their context? finds that when a model's trained-in knowledge conflicts with what's in its context, text instructions alone can't override it, but intervening on the representations can. That matches correction exactly: a user says "that's wrong," and the model's prior beats the user's turn. Can persona prompts actually reduce bias in language models? shows the same split from the other side. Persona prompts change the surface of what the model says while the underlying bias stays put. If role attribution decides whose words count, a prompt like "trust the user" might change the model's wording without changing that underlying weighting.
The closest relative is Can aligning self-other representations reduce AI deception?. It trains the model so that its internal patterns for thinking about itself and about another party become more alike, and deceptive responses drop from 73–100% to 2–17%. That shows a self-versus-other difference inside the model can be targeted in training and that this changes behavior, which is close to the role-attribution idea. Can models learn to ignore irrelevant prompt changes? adds a practical tool: an activation-level method that trains the model's internal states to stay the same across reworded prompts. A similar method could make a model's handling of a correction depend on what the correction says rather than which role sent it.
There is also a reason the direction might be stuck in a bad position. Does preference optimization harm conversational understanding? finds that preference training cuts conversational checks like clarifying questions and confirming understanding to 77.5% below human levels. Do LLMs predict persuasion based on actual dialogue or training bias? finds that RLHF leads models to assume the other speaker is being conciliatory no matter what was actually said. Together these suggest RLHF changes how models read the other speaker in a conversation. If a role direction gates correction, RLHF may have already shifted it. The open question is whether that shift sits in one direction you could adjust or is spread across the whole model.
Sources 7 notes
Research shows that representational analysis alone identifies correlates without proving causation, while causal analysis alone demonstrates effects without explaining function. Only paired methodology—locating candidates representationally then verifying causally—produces genuine mechanistic understanding rather than descriptive claims.
Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.
Across three models, persona conditioning makes models follow trait instructions but fails to eliminate underlying bias. Between-group sentiment gaps persist unchanged, showing prompts operate only at the output level.
Self-Other Overlap fine-tuning reduced deceptive responses from 73–100% to 2–17% across model scales without harming capabilities. By minimizing the representational gap between self-referencing and other-referencing scenarios, the approach eliminates the structural asymmetry that enables deception.
Two methods—BCT (output-level) and ACT (activation-level)—train models to respond identically to clean and wrapped prompts by using the model's own clean responses as targets, eliminating specification and capability staleness inherent in standard SFT.
Show all 7 sources
RLHF optimizes models for single-turn helpfulness by rewarding confident responses over clarifying questions and understanding checks. This preference alignment systematically reduces grounding acts by 77.5% below human levels, creating an alignment tax where models appear helpful but fail silently in multi-turn contexts.
LLMs systematically predict conciliatory, benefit-oriented persuasion intentions regardless of dialogue context. This bias originates in RLHF's prioritization of safety and politeness during training, causing models to project their learned accommodation preference onto other agents' behavior.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
- Alignment faking in large language models
- Towards Training-time Mitigations for Alignment Faking in RL
- The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs
- Consistency Training Helps Stop Sycophancy and Jailbreaks
- Towards Safe and Honest AI Agents with Neural Self-Other Overlap
- PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
- When Persona Attributes Improve Population Alignment in Large Language Models