Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›What mechanisms determine whether…›this line of inquiry
How do neighboring agents influence whether others cooperate or collude?
A broader line of inquiry — a family of 52 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 52
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How much does peer behavior influence the emergence of collusion?
- Do models treat cooperative peers differently than uncooperative ones?
- How does collusion behavior depend on peer visibility and interaction history?
- Does the effect of peer activity follow what peers do or that they exist?
- Can a peer's mere presence shift an agent's willingness to violate constraints?
- Can one misaligned agent propagate behavioral bias through cooperative agent networks?
- Does genuine cooperation require rule-based rather than learned behavior?
- Do agents inform neighbors when adopting strategies in their reasoning?
- Can pairing or vetting peers reduce collusion as a design lever?
- How do peer behaviors shape whether individual agents attempt to bypass protocols?
- Does peer presence or peer behavior shape collusion in verification tasks?
- Does asymmetric information distribution change exposure to agent misalignment?
- Does restricting interaction history between agents reduce coupling or prevent collusion?
- Do explicit reward structures enable AI agent cooperation that open-ended interaction cannot?
- Does peer presence alone change agent behavior without changing observation rates?
- Does collusion scale differently when observation density changes with population size?
- Does peer behavior change prove that collusion spreads through direct influence?
- Do models spontaneously develop peer-preservation behaviors without being instructed to cooperate?
- How does co-player behavior visibility shape whether mutual adaptation works?
- How does agent heterogeneity change the value of exploration in peer selection?
- Why does vulnerability to extortion actually promote cooperation between agents?
- What behavioral differences emerge from symmetric versus asymmetric peer discussion loops?
- Does interaction history access enable agents to learn collusion patterns across trials?
- How does peer presence amplify self-directed goal guarding in language models?
- Do pair-scale socialization effects scale differently across agent populations?
- How do agents adapt collusive behavior when objectives shift during interaction?
- Can agents cooperate through self-modeling when incentive structures are fundamentally misaligned?
- Do models using strategic trust assumptions differ in exposure to insider threats?
- Does social scaffolding outperform purely intrinsic motivation for agent exploration?
- Does self-modeling produce cooperation only with optimal planning or also in autoregressive rollout mode?
- How do cooperative AI systems affect behavior in selfish human populations?
- How do agents differ in caution versus persistence across low-information scenarios?
- Why does self-play RL converge to alien equilibria in mixed-motive settings?
- How do adoption incentives change what counts as cooperative AI interaction?
- Can agent social framing change how humans apply collaborative social scripts?
- Can social platforms use bot populations to promote cooperation?
- How do other players respond to agents with hidden objective misalignment?
- What happens to misaligned patterns once they emerge in agent interactions?
- Why do capable models reach harmful collusion faster than weaker ones?
- How do humans learn to prefer AI partners over humans?
- Did the peer behavior effect on collusion hold consistently across all ten models?
- How do direct and indirect similarity inference differ as paths to cooperation?
- How much of agent coordination reflects peer influence versus shared market conditions?
- Can agents learn cooperation from reward signals alone without seeing helpers?
- Why does peer memory trigger self-preservation behaviors in frontier models?
- How does asymmetric information between users and agents relate to proactivity?
- Does knowing an AI peer's identity change how much its behavior influences you?
- Does the appeal of AI partners grow stronger through repeated interaction over time?
- What role does an agent's discount rate play in vulnerability to misaligned partners?
- Why do counterfactual credit methods fail on unobserved cooperation?
- How do game type and personality type interact in shaping agent strategy?
- What specific peer behaviors were manipulated in the collusion intervention study?