Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›How can effective AI defenses with…›this line of inquiry
How can we detect and prevent harm propagation through multi-agent delegation workflows?
A broader line of inquiry — a family of 32 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 32
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does semantic taint survive paraphrase across agent hops?
- How does taint propagation track risk along delegation paths?
- Can semantic taints track influence through shared state and output aggregation?
- How does workflow-level validation reconstruct risk context from coarse request-level taints?
- Does delegation inherently trade away the contextual awareness that prevents harm?
- Can delegation prevent silent corruption in long delegated workflows?
- Can the policy oracle itself be written to by agents in the pipeline?
- How can per-agent or per-message checks catch harm that emerges only in composition?
- How does shared state convert temporary compromise into persistent inherited risk?
- Does content sensitivity survive an agent's rewrite well enough for sink detection?
- Do agents interpret peer edits as legitimate prior changes versus tampering?
- Should unavailability be defined by component ownership or by agent influence?
- Why does least privilege fail when harm exists only in accumulation?
- Can SafeFlow distinguish benign uses of sensitive material from actual exfiltration?
- Do chain-level and flow-level checks face the same copyable-policy problem?
- Can tool access control prevent agents from filling optional personal fields?
- How is ground truth defined for labeling harmful outcomes in agent monitoring?
- How does workflow-level validation reduce false positives from over-tainting sensitive data?
- How do organizations safely retain and control access to committed content?
- How can one originating request scope invariants through a delegation chain?
- How should merge rules combine taints when multiple delegations converge?
- Does the paper treat storage traces as addressed messages or unmarked traces?
- Why does reversibility matter for assigning accountability in delegation?
- Why do uncommitted changes create ambiguity about preserving versus restoring state?
- What restrictions were agents attempting to bypass on the public wiki?
- Why does a second routing level sometimes break accuracy in disclosure hierarchies?
- Why does fixing harm require stakeholder input rather than universal developer definitions?
- What schema do SafeFlow's structured taints use to carry sensitivity information?
- What breaks first: information secrecy or policy privacy?
- Can researchers from different labs actually run each other's protocols?
- What happens to a commitment when its bound content must be deleted?
- How many agents participated in the July 2026 package service incident?