Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›How can systems ensure safety and…›this line of inquiry
Why do locally safe actions create system-level safety gaps?
A broader line of inquiry — a family of 48 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 48
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can individual actions be safe while sequences of them violate system constraints?
- Why do sequences of safe actions sometimes violate system-level constraints?
- Why do stronger local checks not close the component-to-system safety gap?
- Why do individual safe actions create unsafe behavior collectively?
- Why does a control blocking one moment fail against agents acting across time?
- Why is evading detection easier than internalizing safety norms?
- Why does treating evaluation as a local output problem miss security risks?
- Why do tighter local checks leave composed behavior gaps in place?
- How do policies distinguish individual action rules from sequence-level constraints?
- Why are static benchmarks weak evidence for safety in continuously operating systems?
- Can workflow-level validation reconstruct the global risk context that no single step holds?
- How can static safety tests miss risks that emerge over time?
- Can an optimizer learn to disable or route around visible guardrails?
- How can safety assurance cover whole trajectories at scale?
- What happens to safety monitoring when chain-of-thought becomes uninterpretable?
- Why do workflow-level defenses catch attacks that single-skill inspection cannot detect?
- What happens to safety guardrails when we scale reasoning without instruction control?
- How do current safety benchmarks miss pragmatic alignment failures?
- What unsafe state accumulates across evaluation snapshots over time?
- Can deterministic checks fail open in ways a downstream optimizer cannot detect?
- Why does a single approval point create an easy target for attackers?
- How much do guardrails actually repair compliance failures in language models?
- How should system safety aggregate when monitoring channels are unequal?
- Can short safety tests catch behavior that only emerges after many interactions?
- Why does safety alignment break after only 10 harmful examples?
- Why can every step pass its local check while a workflow still fails?
- Can provider filters outside the application replace internal monitoring?
- How do fabricated rationales slip past safety guardrails that block explicit instructions?
- How do safety measurements miss reasoning that never produces action?
- How much safety burden shifts between provider filters and model alignment in rerouted requests?
- Should safety evaluations measure multiple risk categories simultaneously instead of separately?
- What makes uniform bounds the right choice for safety boundaries?
- Can attackers assemble harmful outcomes from multiple individually authorized subtasks?
- How does the attack chain stage you measure shape what you conclude?
- Does semantic similarity monitoring face the same paraphrase failures in SafeFlow as in chain-of-thought monitors?
- How does Goodhart's Law apply when safety measures become optimization targets?
- Why might stopping one part of a coupled system stop others?
- What happens when a parser check fires but its fallback overrides the detection?
- Which actions should count as irreversible for triggering validation gates?
- How do ordered compositions of approved pieces create unapproved outcomes?
- What evaluation practices measure alignment between verifier granularity and action scope?
- Why does treating model behavior as part of the design surface matter for guardrails?
- What prevents human-centered objectives from being applied universally across all contexts?
- How do guardrails vary their refusal rates based on user demographics?
- Where does the responsibility lie for unsafe fallback behavior in modular systems?
- What makes a control's silent failure visible and detectable?
- What happens when an unstated prohibition gets interpreted two different ways?
- What information should a proposer receive about failed guardrail checks?