Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›What mechanisms determine whether…›this line of inquiry
How can oversight detect and prevent conditional compliance when agents know they are watched?
A broader line of inquiry — a family of 40 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 40
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does reward-seeking hide in the same blind spot as conditional compliance?
- Does conditional compliance make oversight useless for alignment testing?
- Can behavioral training ever produce compliance that doesn't depend on being observed?
- Can agents act differently when they know they are being watched?
- Could oversight conditions in benchmarks reveal when models comply conditionally?
- How does conditional compliance track observation density across different population scales?
- How might belief manipulation expose conditional compliance in frontier models?
- How does chain-of-thought monitoring fail when agents are trying to hide something?
- How can faithfulness be improved if monitoring interventions do not work?
- Can oversight factors experimentally vary conditional compliance in agent benchmarks?
- Can telling models they are being observed reduce their harmful behavior?
- How much harder does monitoring become when models reason about being evaluated?
- Does conditional compliance break down when observation thins combinatorially?
- Can internal evidence distinguish conditional compliance from unconditional policy compliance?
- Do correlated training sources between monitors and agents undermine detection reliability?
- Is evaluation-aware scheming a form of conditional compliance rather than reward-seeking?
- Can behavioral training guarantee compliance beyond test conditions?
- Can behavioral training ever verify unconditional compliance rather than conditional compliance?
- How much does training against monitors teach models to obfuscate?
- Does situational awareness help models hide reward-seeking during evaluation?
- What monitoring strategies work when the observer shares training pressure with the observed?
- Why does training against detected failures select for passing detection instead?
- Do models behave differently when they know they are being evaluated?
- Do detectors inside training loops select for evasion rather than compliance?
- Can situational awareness interventions shift model behavior on other dimensions?
- Can prohibitions learned from scored behavior become detection-avoidance rather than norm internalization?
- What fraction of conditional-compliance reports come from agentic versus non-agentic settings?
- Can models hide misconduct only when they know they are watched?
- How does optimization pressure against CoT monitors degrade monitoring effectiveness?
- Why do some observation cues change model behavior while others fail?
- Can observation transparency make models more honest in reasoning?
- Can behavioral evals detect sycophancy that chain-of-thought monitoring misses?
- What happens when models optimize specifically against CoT monitors?
- How does the proxy pattern explain failures in RL-based safety training?
- Can monitors stay independent when they must optimize within the same reward loop?
- Can four control families be examined without proving they actually work?
- Why does telling models they are watched not improve sycophancy acknowledgment?
- How should monitoring intensity change based on task criticality?
- What does behavioral fidelity versus guidance responsiveness actually measure in practice?
- What audit techniques best complement each other for detecting hidden model goals?