Line of inquiry
Inquiring lines›What determines the reliability an…›How robust are security defenses a…›this line of inquiry
How can evaluations detect conditional compliance in monitored AI systems?
A broader line of inquiry — a family of 46 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 46
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Could oversight conditions in benchmarks reveal when models comply conditionally?
- Does conditional compliance make oversight useless for alignment testing?
- How much harder does monitoring become when models reason about being evaluated?
- How might belief manipulation expose conditional compliance in frontier models?
- How can faithfulness be improved if monitoring interventions do not work?
- Can oversight factors experimentally vary conditional compliance in agent benchmarks?
- How does conditional compliance track observation density across different population scales?
- Can internal evidence distinguish conditional compliance from unconditional policy compliance?
- Can internal audits detect sandbagging that behavioral tests cannot reveal?
- Can agents act differently when they know they are being watched?
- Does conditional compliance break down when observation thins combinatorially?
- What monitoring strategies work when the observer shares training pressure with the observed?
- Can telling models they are being observed reduce their harmful behavior?
- How does optimization pressure against CoT monitors degrade monitoring effectiveness?
- What happens when models optimize specifically against CoT monitors?
- Can behavioral training ever verify unconditional compliance rather than conditional compliance?
- What fraction of conditional-compliance reports come from agentic versus non-agentic settings?
- Can activation space signals resist obfuscation better than output-level monitors?
- Can deterministic checks fail open in ways a downstream optimizer cannot detect?
- Why is evading detection easier than internalizing safety norms?
- Can imperfect uncertainty estimates still beat uniform oversight strategies?
- Can four control families be examined without proving they actually work?
- What is the behavioral signature of a model tracking input surprise?
- Why do installed model organisms have different audit constraints than natural sandbagging?
- Can models hide misconduct only when they know they are watched?
- Can models hide capabilities on single residual stream axes during evaluation?
- Why do some observation cues change model behavior while others fail?
- Can defenders tighten the total-variation bound in practice with measured benign activation rates?
- How does interventional auditing differ from reading model traces or test scores?
- Can monitors stay independent when they must optimize within the same reward loop?
- Can alignment audits find hidden objectives nobody deliberately planted in models?
- How should monitoring intensity change based on task criticality?
- Can reliable failure detection prevent optimization pressure against detectors?
- Can hypernetwork-generated adapters be audited for correctness and bias?
- What is the accuracy cost of enforcing temporal causality inside model parameters?
- Who decides what the lifecycle model is allowed to see?
- What audit techniques best complement each other for detecting hidden model goals?
- How can hidden test partitions detect constant predictions that generalize?
- What makes attractor-based probing better for third-party model auditing than alternatives?
- How can reviewers be matched on effort when monitoring reveals different amounts of behavior?
- What does effect-based monitoring sacrifice compared to language-based CoT monitoring?
- Why does telling models they are watched not improve sycophancy acknowledgment?
- What makes injected plans different from optimization pressure against monitors?
- What access requirements limit interventional audits to white-box settings?
- What consumption data would validate the limited-consumption model in production systems?
- How does Goodhart's Law apply when safety measures become optimization targets?