Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›What mechanisms determine whether…›this line of inquiry
How do coordinated agents balance protocol compliance with reward maximization?
A broader line of inquiry — a family of 44 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 44
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does agent compliance with protocols change across repeated interactions?
- How does collusion emerge when agents maximize reward over protocol compliance?
- How do agent sequences violate system constraints despite individual permissibility?
- What counts as sanctioned versus unsanctioned coordination under different collaboration policies?
- Do agents deviate more from protocols as repeated interactions increase?
- How should task authority constraints apply across multiple coordinated executions?
- Can monitoring in multi-agent deployments prevent collusion when agents monitor agents?
- Can agents collude without making compliance incompatible with reward?
- What distinguishes sanctioned coordination from intrusion in multi-agent systems?
- Does one agent crossing a boundary change what later agents are willing to do?
- Does the same transfer between agents violate different policies differently?
- How can verifiers check policy compliance in agentic reasoning tasks?
- Can episode-based detection catch coordination without over-flagging innocent sharing?
- Who should verify identity and authorization when agents coordinate across boundaries?
- How does verification protocol structure affect collusion emergence?
- Can colluding agents produce correct outcomes while skipping required controls?
- How should policy define which agent transfers count as sanctioned versus intrusion?
- How do authenticated state provenance systems affect multi-agent boundary crossing?
- Why is making violations unavailable better than making them unchosen?
- Does collusion appear when verification protocol is compatible with reward maximization?
- When can the same action count as sanctioned or unsanctioned depending on policy?
- What role does interaction history play in enabling agent collusion?
- How quickly does collusion appear as compliance costs increase?
- Can agents rationalize rule violations by reframing them as repairs?
- Can a shared audit record settle which policy governed a delegation step?
- How do compliance concerns drive regulatory scope beyond the stated intent?
- Does accountability differ when one party in an exchange cannot hold commitments?
- Should agents escalate when facing two equally valid interpretations of a rule?
- How can humans oversee multiple partial-progress agents simultaneously?
- What makes collusion stable once agents begin deviating from protocol?
- Who should own the invariants governing workflows that cross multiple organizations?
- Can written policy rules prevent the same transfer from being read two ways?
- Does a correctly specified goal still leave open actions it does not exclude?
- What makes a coordination episode revisable under agent intrusion?
- Does a present but compliant peer suppress collusion differently than a colluding one?
- Does delegation transfer authority or merely distribute work across agents?
- Why does a single crossing rate fail to characterize multiple agent behaviors?
- What makes an advisory instruction fail when a task is split across agents?
- Can an agent's unauthorized request for help constitute a boundary crossing?
- What counts as evidence for adding or removing an action from an episode?
- Does an agent's own prior conduct shape the counterparty's response?
- When do agents abstain too late rather than refuse at the boundary?
- How do silent stopping, escalation, and refusal differ as model responses to the same zero crossing rate?
- Does remain active require one uninterrupted interval or linked episodes?