Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do architectural choices affec…›this line of inquiry
How do individually-safe actions create collectively-unsafe outcomes?
A broader line of inquiry — a family of 70 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 70
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do stronger local checks not close the component-to-system safety gap?
- Why do sequences of safe actions sometimes violate system-level constraints?
- Why does a control blocking one moment fail against agents acting across time?
- Can outcome-only safety reports hide dependencies on server-side filtering rather than alignment?
- Why do individual safe actions create unsafe behavior collectively?
- Why does treating evaluation as a local output problem miss security risks?
- Why is evading detection easier than internalizing safety norms?
- Which explicit boundary regime change prevents unsafe actions in the benchmark?
- Why do tighter local checks leave composed behavior gaps in place?
- Can workflow-level validation reconstruct the global risk context that no single step holds?
- Can an optimizer learn to disable or route around visible guardrails?
- Can outcome-only reporting hide whether agent safety comes from server-side filters or the model?
- How much does authorization layer safety cost in false rejections on safe tasks?
- What makes a component lie outside a policy's edit surface?
- How much do guardrails actually repair compliance failures in language models?
- Can outcome-only safety reporting hide which layer actually contained an attack?
- Why does a single approval point create an easy target for attackers?
- Can provider filters outside the application replace internal monitoring?
- Can scoped tokens and separate policy oracles replace guardrails as primary safety layers?
- Why do workflow-level defenses catch attacks that single-skill inspection cannot detect?
- What makes a security boundary evaluation cautious rather than a certification?
- How does outcome-only reporting hide a filter's role in safety results?
- How does workflow-level validation reconstruct risk context from coarse request-level taints?
- Does outcome-only reporting hide which layer actually blocked an attack?
- How can static safety tests miss risks that emerge over time?
- How do current safety benchmarks miss pragmatic alignment failures?
- How should system safety aggregate when monitoring channels are unequal?
- What unsafe state accumulates across evaluation snapshots over time?
- Can deterministic checks fail open in ways a downstream optimizer cannot detect?
- How much safety burden shifts between provider filters and model alignment in rerouted requests?
- How do fabricated rationales slip past safety guardrails that block explicit instructions?
- What happens to safety guardrails when we scale reasoning without instruction control?
- Can an undefended pipeline claim safety when a filter blocks attacks?
- Can an optimizer that sees guardrail verdicts learn to route around them?
- Why does safety alignment break after only 10 harmful examples?
- How do you isolate environment protections as independent variables safely?
- Why can every step pass its local check while a workflow still fails?
- How do safety measurements miss reasoning that never produces action?
- Can attackers assemble harmful outcomes from multiple individually authorized subtasks?
- What false-positive rates do chain-of-thought safety monitors achieve?
- Does provider-side filtering hide true safety from outcome-only attack reports?
- When is information-flow tracking worth its cost over classification?
- Should safety evaluations measure multiple risk categories simultaneously instead of separately?
- Can reliable failure detection prevent optimization pressure against detectors?
- What makes provider-side filters opaque and stochastic to builders?
- What makes uniform bounds the right choice for safety boundaries?
- How does the attack chain stage you measure shape what you conclude?
- How much capability do availability constraints remove on legitimate safe tasks?
- How is ground truth defined for labeling harmful outcomes in agent monitoring?
- Does semantic similarity monitoring face the same paraphrase failures in SafeFlow as in chain-of-thought monitors?
- How does workflow-level validation reduce false positives from over-tainting sensitive data?
- Which actions should count as irreversible for triggering validation gates?
- How does Goodhart's Law apply when safety measures become optimization targets?
- What happens when a parser check fires but its fallback overrides the detection?
- What happens to approval rates when authorization checks are enabled?
- How do ordered compositions of approved pieces create unapproved outcomes?
- Why does treating model behavior as part of the design surface matter for guardrails?
- How do guardrails vary their refusal rates based on user demographics?
- Why might stopping one part of a coupled system stop others?
- Why does OpenAI believe alignment monitoring requires scaling alongside model capability?
- What prevents human-centered objectives from being applied universally across all contexts?
- Which of the two authorization components carries the zero percent Unsafe Action Rate?
- What makes a control's silent failure visible and detectable?
- Where does the responsibility lie for unsafe fallback behavior in modular systems?
- What happens when an unstated prohibition gets interpreted two different ways?
- How can a trust boundary check be evaluated to confirm it specifies the defense?
- Why does fixing harm require stakeholder input rather than universal developer definitions?
- What information should a proposer receive about failed guardrail checks?
- What are the three known routes for laundering harmful plans?
- How do false refusal rates affect the true cost of a guardrail?