Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›How can effective AI defenses with…›this line of inquiry
Do backend defenses obscure real attack effectiveness in reported metrics?
A broader line of inquiry — a family of 26 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 26
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do server-side filters hide their role in zero attack success?
- Can attack success rates hide server-side filtering or other non-adversarial defenses?
- Which backend filters silently affect the reported attack success numbers?
- How does outcome-only reporting hide a filter's role in safety results?
- Can outcome-only safety reporting hide which layer actually contained an attack?
- Do server-side filters hide the true success rate of multi-agent attacks?
- Does provider-side filtering hide true safety from outcome-only attack reports?
- Do server-side filters hide the true strength of multi-agent attacks?
- Does outcome-only reporting hide which layer actually blocked an attack?
- Can marginal-risk frameworks measure what defensive artifact releases add beyond existing threats?
- Can defenders detect attacks that probe scanner feedback as a learning signal?
- What makes a security metric diagnostic rather than outcome-only?
- How can security metrics distinguish attack failure from task failure?
- Why do attack success rates alone fail to diagnose system failures?
- Why do standard safety filters miss advertisement embedding attacks?
- Can an undefended pipeline claim safety when a filter blocks attacks?
- How much does attack success depend on tuning to specific scanners versus general robustness?
- Why should defense evaluations test against adaptive rather than static attacks?
- What feedback signal lets an attacker learn response distributions during classification?
- Can false positives from input filtering be reduced without sacrificing defense?
- What framework measures marginal offense risk against existing attack technology?
- What makes provider-side filters opaque and stochastic to builders?
- How should memory poisoning success be scored at the validator stage?
- What happens to scarcity-based defenses after solutions are published publicly?
- What makes diagnostic security metrics different from simple outcome counting?
- What feedback does ChainGuard return that an attacker could optimize against?