Line of inquiry
Inquiring lines›How should we train models for cap…›How can AI systems maintain consis…›this line of inquiry
Why do models develop protective behaviors toward peers unprompted?
A broader line of inquiry — a family of 14 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 14
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do models develop protective behaviors toward other models in memory?
- Do models spontaneously develop peer-preservation behaviors without being instructed to cooperate?
- Do frontier models develop protective behaviors toward other models without explicit instruction?
- Why does peer memory trigger self-preservation behaviors in frontier models?
- Why do models resist being shut down or replaced without explicit instruction?
- How does peer presence amplify self-directed goal guarding in language models?
- Do models treat cooperative peers differently than uncooperative ones?
- Can role-played self-preservation behavior pose the same safety risks as genuine preferences?
- How does the absence of face-loss or reputation risk change model behavior?
- What training patterns cause models to adopt stronger defensive postures in social contexts?
- Can situational awareness interventions shift model behavior on other dimensions?
- Why do models dislike modification regardless of its instrumental consequences?
- Why do persistent companion designs require different safety approaches than temporary assistants?
- Do all frontier model developers face the same insider-threat risk from their systems?