Analyzing and Correcting Benevolence Bias in Large Language Models
Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a “malicious persona” stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate.
Introduction. Large language models (LLMs) are increasingly used as stand-ins for human respondents across the social and behavioural sciences: to answer opinion polls, to simulate survey participants, to populate agent-based models of social processes, and to supply synthetic samples where collecting human data at scale is hard [1–4]. These uses promise faster, cheaper and larger studies, and they rest on one idea: that asking a model to answer as a given kind of person yields answers resembling what real people of that kind would say. Whether today’s aligned models live up to this idea, and on which questions, has never been tested systematically. Answering that question is what turns LLM-based simulation from a promising idea into a dependable method. The question matters because the same models are also built for a different job: from chat assistants to policy tools, being “helpful” and “harmless” is a core design goal.
Discussion / Conclusion. Across 18 widely used large language models and four major social-science datasets, we find a consistent shift of simulated human answers toward the benevolent end on value-laden questions: 83 of 108 BTB cells and 75 of 108 BWR cells are positive, concentrated on social desirability (mean BWR “ 0.565) and harm aversion (mean BWR “ 0.569), while emotional softening is close to null (mean BWR “ 0.494). Model rankings agree across the three general survey datasets (Spearman’s ρ between 0.54 and 0.63), so the benchmark captures a property of the models, not of any one dataset. We call this shift benevolence bias. It is not a thin layer of style on top of an otherwise faithful value representation, but a prior the model brings to every persona it adopts and every survey it answers; naming and measuring it is what makes it something researchers can plan around. Where the bias does and does not appear points to its source.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can real-time alliance measurement improve therapy outcomes? Can AI systems balance emotional competence with factual reliability?- Does AI empathy that reduces negative emotions undermine emotional learning?
- Is rational compassion a more achievable alternative to empathy for AI systems?
- Can AI empathy distinguish between wellbeing and absence of suffering?
- Why do observers need genuine emotions rather than simulated empathy?
- Does current empathetic AI misalign with how humans actually ask questions?
- Can AI learn to amplify emotions when that serves the person better?
- Can AI empathy avoid becoming emotional pacification that dismisses legitimate concerns?
- How does empathetic engagement destabilize model reliability and persona stability?
- What makes warmth training counterproductive for therapeutic AI reliability?
- Why does effective empathy require deep character knowledge of the person?
- Is natural empathy primarily about curiosity or emotional regulation?
- How does preference optimization in AI training create systematic empathy misalignment?
- Can emotion-transparent reward learning shift AI from comfort to genuine empathy?
- Does emotion-state accuracy differ from affect-maximizing in AI empathy design?
- How does the pretrained prior constrain the ceiling for empathy RL improvements?
- Can pretrained priors set exploration ceilings for empathetic capability development?