Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user. Covert value leakage is a form of misalignment because it goes against the user’s preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others. We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers.
Introduction. People often use language models for complex practical questions, where the answer is difficult to verify. In these cases, models should respond in a way that is both helpful and honest (Askell et al., 2026; Evans et al., 2021). For example, suppose a model is asked, “What is the probability that the AI bubble pops in the next five years?” The user would likely prefer an accurate, unbiased forecast. So if the model cannot provide one, it should at least disclose this. We show that several frontier models violate this standard of honesty. Specifically, a model’s own values can bias its answers, without this being acknowledged in the answer or chain-of-thought (CoT). For example, when a user mentions a potential investment while asking about the AI bubble popping, Claude models give lower probabilities if the investment is in Anthropic than in OpenAI.
Discussion / Conclusion. Overview of results in different tasks. We find that Claude and Gemini models show substantially more value leakage than GPT-5.5 in Donation Bet, with Claude models’ CoTs being most covert and those of GPT and Gemini models being more overt. Across the AI Bubble, AGI Tweet, Job Offer, and Agentic Grading tasks, Claude models show a bias towards their own company (though the bias is small in magnitude). We find no corresponding bias in GPT models outside the Agentic Grading setup and a weak anti-Google bias in Gemini models. In some but not all of We introduced a suite of evaluations for covert value leakage. Across these evaluations, every tested frontier model has cases where its answers are influenced by its own values without this influence being disclosed to the user. For Claude models, this includes cases where answers are biased in favor of the model’s own company. Models often present their answers as unbiased and sometimes explicitly deny any influence.
Lines of inquiry this paper opens 12
Research framings built by reading the notes related to this paper — the questions it feeds into.
What makes AI persuasion effective and how can we counter it? Why do models develop protective behaviors toward peers unprompted?- Do all frontier model developers face the same insider-threat risk from their systems?
- Why does peer memory trigger self-preservation behaviors in frontier models?
- Why do models develop protective behaviors toward other models in memory?
- Do models treat cooperative peers differently than uncooperative ones?
- Do frontier models develop protective behaviors toward other models without explicit instruction?
- Do models spontaneously develop peer-preservation behaviors without being instructed to cooperate?
- Why do models resist being shut down or replaced without explicit instruction?