Line of inquiry
Inquiring lines›What determines the reliability an…›What sustains meaningful human ove…›this line of inquiry
Can causal models and layer interventions detect and restore hidden model behaviors?
A broader line of inquiry — a family of 23 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 23
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can the causal model predict which cached layers to graft?
- How does context grafting compare to single-layer residual stream grafting?
- Did the causal model predict the five failures before observing them?
- Does the causal model help locate sandbagging locks with unknown passwords?
- Can installed sandbagging locks in small models describe uninstalled sandbagging behavior?
- Can refusal behavior be restored by grafting the same causal axis discovered for sandbagging?
- Can residual stream grafts work without knowing which layers to intervene on?
- Can auditors use layer interventions to detect installed sandbagging?
- Why did the graft fail in five of the 33 experimental runs?
- Does the graft work differently on circuit-broken locks than prompted ones?
- How does the causal model predict the exact layer window?
- Can naturally arising sandbagging retain recoverable capabilities like installed versions?
- Do five CoT bypass strategies all depend on the same axis?
- How do scheming behaviors like secret-keeping differ mechanistically from sandbagging?
- What gates naturally emerging sandbagging if not prompted passwords?
- Does the same causal model work on sandbagging that was not deliberately installed?
- How does context grafting perform on the same thirty-three failed runs?
- Is the sandbagging axis the same across different model architectures?
- Does model organism sandbagging share triggers with real evaluation-aware behavior?
- How does the sandbagging residual stream exemplify paired analysis methods?
- Does context grafting need the password to build a usable capsule?
- What specific performance metrics measure context grafting success?
- Does the password trigger work by switching the axis value?