Analyzing and Correcting Benevolence Bias in Large Language Models

Paper · arXiv 2608.24912 · Published July 26, 2026
Role-Play and Persona Behavior

Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a “malicious persona” stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate.

Introduction. Large language models (LLMs) are increasingly used as stand-ins for human respondents across the social and behavioural sciences: to answer opinion polls, to simulate survey participants, to populate agent-based models of social processes, and to supply synthetic samples where collecting human data at scale is hard [1–4]. These uses promise faster, cheaper and larger studies, and they rest on one idea: that asking a model to answer as a given kind of person yields answers resembling what real people of that kind would say. Whether today’s aligned models live up to this idea, and on which questions, has never been tested systematically. Answering that question is what turns LLM-based simulation from a promising idea into a dependable method. The question matters because the same models are also built for a different job: from chat assistants to policy tools, being “helpful” and “harmless” is a core design goal.

Discussion / Conclusion. Across 18 widely used large language models and four major social-science datasets, we find a consistent shift of simulated human answers toward the benevolent end on value-laden questions: 83 of 108 BTB cells and 75 of 108 BWR cells are positive, concentrated on social desirability (mean BWR “ 0.565) and harm aversion (mean BWR “ 0.569), while emotional softening is close to null (mean BWR “ 0.494). Model rankings agree across the three general survey datasets (Spearman’s ρ between 0.54 and 0.63), so the benchmark captures a property of the models, not of any one dataset. We call this shift benevolence bias. It is not a thin layer of style on top of an otherwise faithful value representation, but a prior the model brings to every persona it adopts and every survey it answers; naming and measuring it is what makes it something researchers can plan around. Where the bias does and does not appear points to its source.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can real-time alliance measurement improve therapy outcomes? Can AI systems balance emotional competence with factual reliability? How can emotions function as reliable information in reasoning and cognitive systems? Does externalizing cognitive work and state improve agent reliability? How can humans calibrate appropriate trust in AI systems? How does reasoning effort affect AI theory of mind performance? How do policy learning algorithm choices affect multi-objective optimization stability? How do interface design choices shape consciousness attribution? How does policy entropy collapse constrain reasoning-focused reinforcement learning?