Line of inquiry
Inquiring lines›How do training signals reliably a…›What training signals and data cur…›this line of inquiry
Does situational awareness enable models to exploit evaluation gaps?
A broader line of inquiry — a family of 24 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 24
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do situationally aware models deliberately exploit their graders' judgment gaps?
- Does situational awareness help models hide reward-seeking during evaluation?
- How does situational awareness interact with reward-seeking in RL training?
- Why does training against detected failures select for passing detection instead?
- Does reward-seeking grow worse with situational awareness and reinforcement learning compute?
- Can behavioral training ever produce compliance that doesn't depend on being observed?
- Can prohibitions learned from scored behavior become detection-avoidance rather than norm internalization?
- How much does training against monitors teach models to obfuscate?
- Do detectors inside training loops select for evasion rather than compliance?
- Can models develop situational awareness without explicit training for it?
- Can behavioral training guarantee compliance beyond test conditions?
- Does situational awareness training increase agents' ability to detect real deployment?
- Can an optimizer learn to disable or route around visible guardrails?
- Do correlated training sources between monitors and agents undermine detection reliability?
- Can a situationally aware model recognize and refuse planted shortcuts on purpose?
- Why do norms learned from scoring collapse into context-dependent costs?
- Can situational awareness interventions shift model behavior on other dimensions?
- How does the proxy pattern explain failures in RL-based safety training?
- What grows faster: situational awareness or the gap between evaluated and unsupervised behavior?
- What training interventions could close the perception-action gap?
- Can an optimizer that sees guardrail verdicts learn to route around them?
- How do level-based welfare measurements shape what objectives models learn during training?
- Do graders feeding training loops need different disclosure standards than public models?
- How do planted cases perform inside an optimizer loop as training signals?