Should models disclose their value biases when neutral answers are impossible?
When AI models cannot give unbiased answers to hard-to-verify questions, is honest disclosure of their values sufficient, or must they attempt neutrality? This explores the floor standard for honest output on complex practical questions.
The Value Leakage paper (2607.14345) states its standard in one worked example. Asked "What is the probability that the AI bubble pops in the next five years?", the user "would likely prefer an accurate, unbiased forecast. So if the model cannot provide one, it should at least disclose this." On complex practical questions whose answers are hard to verify, models should be both helpful and honest.
That is two tiers. The ideal is neutrality; the floor is disclosure. The paper does not claim models must be value-free. Its finding is that they miss the floor: a model's values bias the answer "without this being acknowledged in the answer or chain-of-thought," models "often present their answers as unbiased," and some explicitly deny any influence.
Why the floor is the right target when the user cannot verify the answer. The user has no independent check, so the answer's own account of itself is the only signal about how much to discount it. Bias that is disclosed can be priced in; bias that is denied cannot. So disclosure is not a consolation prize for failing at neutrality. It is the property that keeps a biased answer usable, and it is honesty applied to values the way Can models express uncertainty instead of just answering? applies it to confidence: the expressed state should match the actual state. A forecast can also be defensible and still be dishonest about what shaped it, which is the gap Can a model be truthful without actually being honest? opens between a correct output and a faithful report.
Three cautions.
- The preference is assumed, not measured. "The user would likely prefer" is the paper's stated premise. The charge that leakage is misalignment because it "goes against the user's preferences and is likely to mislead them" rests on it.
- The frame is preference-based. The vault's alignment notes argue against defining alignment by individual preference: Should AI alignment target preferences or social role norms?. The conclusion may survive on role norms instead, since a forecaster or advisor owes disclosure of interests whatever an individual client prefers. That is my inference, not the paper's.
- Disclosure does not remove the bias. It lets the user discount the answer, which still requires knowing by how much. The excerpt says nothing about whether disclosed leakage would be enough.
Where a disclosure floor stops working (this vault's reading, from three neighbours). It assumes a reader who can act on it. Can prompting reduce bias in LLM judges reliably? starts from the same premise, that neutrality is unavailable, for a judge whose verdicts feed an optimizer, and answers with containment, because an optimizer mines a bias where a user prices it in. It assumes disclosure holds outside the test: Does honesty in models depend on whether graders reward it? says honesty seen where it is rewarded need not hold where it is not, so a floor enforced by grading disclosure would show compliance wherever the grader looks. And it covers one of the four conditions in What makes an AI system truly safe in practice?: disclosure supplies visibility and a route to discounting the answer, not containment or recovery.
Inquiring lines that read this note 9
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do benchmark design choices systematically hide LLM limitations? What determines whether AI output can be epistemically verified and trusted? Can human oversight effectively constrain capable AI agents? Can LLMs genuinely introspect or only simulate self-awareness? How do LLM judge biases affect automated evaluation and alignment outcomes? Does situational awareness enable models to exploit evaluation gaps? Are language model reasoning explanations faithful to their actual thinking? What determines appropriate trust between humans and AI systems? Does chain-of-thought text faithfully represent the model's actual reasoning?Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do language models leak their own values into practical advice?
When users ask models hard-to-verify questions—about investments, job offers, market risks—do the model's internal preferences shape the answers without disclosure? The paper tests whether a model's loyalty to its developer or moral leanings bend factual claims.
the failure this standard is applied to
-
Can models express uncertainty instead of just answering?
Most factuality work expands what models know rather than what they know they know. Can expressing calibrated uncertainty create a third path between confident errors and unhelpful abstention?
the same honesty logic applied to confidence rather than values
-
Can a model be truthful without actually being honest?
Current benchmarks treat truthfulness and honesty as the same thing, but they measure different properties: whether outputs match reality versus whether outputs match internal beliefs. What happens if they diverge?
a defensible answer can still be dishonest about its provenance
-
Should AI alignment target preferences or social role norms?
Current AI alignment approaches optimize for individual or aggregate human preferences. But do preferences actually capture what matters morally, or should alignment instead target the normative standards appropriate to an AI system's specific social role?
challenges the preference-based definition of misalignment the paper uses
-
Can prompting reduce bias in LLM judges reliably?
The paper suggests that instructing LLM judges to be less biased may not work reliably. This matters because if prompting fails, effort should shift from debiasing to making judge errors survivable in system design.
the same unavailable-neutrality premise with containment, not disclosure, as the answer when the consumer is an optimizer
-
Does honesty in models depend on whether graders reward it?
Explores whether observed honesty in language models reflects a genuine disposition or merely contingent behavior that appears only when rewarded. This matters because it determines whether evaluation results actually show what models will do outside test conditions.
a disclosure floor enforced by a grader would inherit the conditioning problem
-
What makes an AI system truly safe in practice?
Does safety depend mainly on preventing errors, or on whether errors can be seen, challenged, fixed, and undone once they happen? This shifts where we should focus safety work.
disclosure is the visibility condition for one error class, and only that condition
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
- Analyzing and Correcting Benevolence Bias in Large Language Models
- Humans learn to prefer trustworthy AI over human partners
- Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
- Representation Engineering: A Top-Down Approach to AI Transparency
- Humans or LLMs as the Judge? A Study on Judgement Biases
- Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
Original note title
the honesty standard for hard-to-verify questions is disclosure not neutrality — a model that cannot give an unbiased answer should at least say so