Interpreting and Steering LLM Agents for Social Simulations
Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions.
Introduction. Large Language Models (LLMs) mimic humans on multiple dimensions: they understand natural Having said that, LLM-based simulations might not always faithfully match human behavior Here, we develop a framework to compare alternate methods for looking into the LLM black-box
Discussion / Conclusion. Future work might also explore hybrid approaches, such as multi-probe steering or low-rank
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can LLM user simulators model realistic goal-driven conversation?- Do emotion-driven actions in agent simulators capture genuine belief revision or just reactive behavior?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- Do realistic LLM behaviors require simulating human thought or just behavior?
- Why does LLM simulation elicit information that direct elicitation cannot?
- Can distributional views explain when an LLM appears to change its mind?
- How do theory of mind and empathy differ in LLM simulation?
- Why do users attribute beliefs to LLMs despite uncertainty about their minds?
- Can agents revise their beliefs predictably when presented with interventions?
- What cognitive capabilities do agents need to internalize social feedback?
- What cognitive structures do realistic belief models need to include?
- Can belief networks from interviews simulate how people change their minds?
- Can causal belief networks extracted from interviews predict how people respond to policy changes?
- How does causal structure avoid behaviorist limitations in LLM social simulation?