Interpreting and Steering LLM Agents for Social Simulations

Paper · arXiv 2609.16436 · Published September 14, 2026
Dialog Topics and Modeling

Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions.

Introduction. Large Language Models (LLMs) mimic humans on multiple dimensions: they understand natural Having said that, LLM-based simulations might not always faithfully match human behavior Here, we develop a framework to compare alternate methods for looking into the LLM black-box

Discussion / Conclusion. Future work might also explore hybrid approaches, such as multi-probe steering or low-rank

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can LLM user simulators model realistic goal-driven conversation? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? What makes AI persuasion effective and how can we counter it? Is model self-awareness based on genuine introspection or pattern matching? Can LLM personas constitute genuine psychology or remain linguistic role-play? Why do multi-turn conversations degrade AI intent and coherence? How does rhetorical adaptation affect LLM persuasion and detectability? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? What prevents language models from reliably adopting diverse personas? Does RLHF training sacrifice accuracy and grounding for user agreement? Why do reward structures fail to shape long-term agent learning? How should models express uncertainty rather than forced confident answers? How do LLMs distinguish causal reasoning from temporal and semantic associations? Do language models develop causal world models or rely on statistical patterns? How do language models inherit human biases from training data?