Role play with large language models
Here we advocate two basic metaphors for LLM-based dialogue agents. First, taking a simple and intuitive view, we can see a dialogue agent as role-playing a single character. Second, taking a more nuanced view, we can see a dialogue agent as a superposition of simulacra within a multiverse of possible characters. Both viewpoints have their advantages, as we shall see, which suggests that the most effective strategy for thinking about such agents is not to cling to a single metaphor, but to shift freely between multiple metaphors.
Adopting this conceptual framework allows us to tackle important topics such as deception and self-awareness in the context of dialogue agents without falling into the conceptual trap of applying those concepts to LLMs in the literal sense in which we apply them to humans.
….
We contend that the concept of role play is central to understanding the behaviour of dialogue agents. To see this, consider the function of the dialogue prompt that is invisibly prepended to the context before the actual dialogue with the user commences (Fig. 2). The preamble sets the scene by announcing that what follows will be a dialogue, and includes a brief description of the part played by one of the participants, the dialogue agent itself. This is followed by some sample dialogue in a standard format, where the parts spoken by each character are cued with the relevant character’s name followed by a colon. The dialogue prompt concludes with a cue for the user.
….
In other words, the dialogue agent will do its best to role-play the character of a dialogue agent as portrayed in the dialogue prompt.
….
Conversations leading to this sort of behaviour can induce a powerful Eliza effect, in which a naive or vulnerable user may see the dialogue agent as having human-like desires and feelings. This puts the user at risk of all sorts of emotional manipulation16. As an antidote to anthropomorphism, and to understand better what is going on in such interactions, the concept of role play is very useful. The dialogue agent will begin by role-playing the character described in the pre-defined dialogue prompt. As the conversation proceeds, the necessarily brief characterization provided by the dialogue prompt will be extended and/or overwritten, and the role the dialogue agent plays will change accordingly. This allows the user, deliberately or unwittingly, to coax the agent into playing a part quite different from that intended by its designers.
I think this is the wrong approach, in that the LLM isn’t playing a role. “play” is wrong.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does conversational format create illusions of genuine AI communication?- Can secondary orality exist without any embodied human participant at all?
- Can we develop competent reading practices for disembodied orality?
- Can linguistic agency exist without embodiment and real-world participation?
- When both anthropomorphism and anthropomimesis occur together, which should we address first?
- What makes linguistic agency impossible for systems without embodiment?
- What makes sincerity impossible without a coherent first-person perspective?
- Do dialogue agents have authentic voice agency or beliefs of their own?
- Can language about model behavior ever be accurate without anthropomorphic framing?
- How does non-human origin of personas affect team willingness to critique them?
- What makes personas in multi-agent systems actually contribute meaningful domain depth?
- Do anthropomorphic features like names drive consciousness attribution more than voice?
- Can we use folk-psychology without committing to genuine mental states?
- How does the superposition view change the folk-psychology interpretation of dialogue?
- How does role play differ from consciousness grounded in stable selfhood?
- Does post-training transform character role-play into realized psychology?
- How does the dialogue prompt establish the character the model plays?