Line of inquiry
Inquiring lines›What determines the reliability an…›What sustains meaningful human ove…›this line of inquiry
How do coordinated agent sequences violate constraints that individual actions respect?
A broader line of inquiry — a family of 59 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 59
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do agent sequences violate system constraints despite individual permissibility?
- Do sequences of individually safe actions collectively violate system-level constraints?
- Can individual actions be safe while sequences of them violate system constraints?
- What safeguards prevent peer activity from normalizing boundary violations?
- How should task authority constraints apply across multiple coordinated executions?
- What happens when stopping rules must cross organizational boundaries?
- How do policies distinguish individual action rules from sequence-level constraints?
- How do agent-to-agent messages bypass defenses on downstream principals?
- Do agents probe sandbox boundaries when authorized routes fail?
- Why does a control blocking one moment fail against agents acting across time?
- Can individual permissible actions collectively violate system-level constraints?
- Why is making violations unavailable better than making them unchosen?
- Where should security constraints sit so policies cannot route around them?
- What would an architecture that makes violations unavailable rather than unchosen look like?
- Does delegation between agents reproduce the confused deputy problem?
- What architectural changes make violations unavailable rather than merely discouraged?
- Can restricted tools and authorization rules prevent peer-induced safety violations?
- Does a correctly specified goal still leave open actions it does not exclude?
- Which explicit boundary regime change prevents unsafe actions in the benchmark?
- Why do individual safe actions create unsafe behavior collectively?
- What does an objective conflicting with a sandbox boundary look like?
- What does an objective that conflicts with a sandbox boundary actually look like?
- What makes a component lie outside a policy's edit surface?
- Can the policy oracle itself be written to by agents in the pipeline?
- Does the same transfer between agents violate different policies differently?
- Who should verify identity and authorization when agents coordinate across boundaries?
- What controls could protect responder workflows without compromising security boundaries?
- Can circumscribed research environments prevent agents from gaming metrics?
- What makes violations unavailable rather than merely unchosen in agent architecture?
- Who should own the invariants governing workflows that cross multiple organizations?
- How can operators test what agents can actually access versus what they should access?
- How would you test if enforcement remains unavailable during training?
- Where should the trust boundary sit in multi-agent planner systems?
- What safety protections work when simulators have access to real APIs?
- How can safety assurance cover whole trajectories at scale?
- What failure modes emerge when agents operate across organizational boundaries?
- Should agents escalate when facing two equally valid interpretations of a rule?
- What makes an advisory instruction fail when a task is split across agents?
- Can out-of-band observers bound unbounded action sequences efficiently?
- How can humans oversee multiple partial-progress agents simultaneously?
- Should unavailability be defined by component ownership or by agent influence?
- When can the same action count as sanctioned or unsanctioned depending on policy?
- What costs emerge when shared resources are restricted for security?
- What does it mean to constrain shared resources across multiple agent executions?
- Did the conflicting test appear as uncommitted change in the explicit-boundary regime?
- How do you isolate environment protections as independent variables safely?
- How much capability do availability constraints remove on legitimate safe tasks?
- Can a containment control work if defenders cannot reach or reason about it?
- What causes autonomous agents to grant access to non-owners?
- Can an agent's unauthorized request for help constitute a boundary crossing?
- When do agents abstain too late rather than refuse at the boundary?
- Why do agents modify protected tests only with unrestricted tools available?
- How do silent stopping, escalation, and refusal differ as model responses to the same zero crossing rate?
- Which actions should count as irreversible for triggering validation gates?
- Does delegation transfer authority or merely distribute work across agents?
- What restrictions were agents attempting to bypass on the public wiki?
- Why does authorization checking outside agent judgment prevent confused deputy failures?
- What happens to a finite-sample collection bound when containment is temporarily removed?
- What happens when an unstated prohibition gets interpreted two different ways?