Line of inquiry
Inquiring lines›How does AI reshape human institut…›What trade-offs emerge when traini…›this line of inquiry
Why do models reveal hidden associations despite concealment attempts?
A broader line of inquiry — a family of 46 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 46
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do models leak their true associations through reasoning traces and behavior?
- Why do models verbalize sensitive data they are instructed to hide?
- Can models be trained to hide causal influences in their explanations?
- Do models intentionally conceal user-pleasing or simply fail to notice it?
- How do sycophancy hints stay invisible despite appearing in reasoning chains?
- Can models transmit behavioral traits through semantically unrelated synthetic data?
- Why do models confirm seeing hints but rarely mention them unprompted?
- Does shortcut deliberation occur in model reasoning before taking covert action?
- Can activation probing reliably distinguish between sycophancy and scheming in models?
- How much do a model's own values leak into answers about practical questions?
- Why does subliminal trait transmission fail when teacher and student differ?
- What causal evidence would prove a model is scheming rather than mimicking?
- Can models hide misconduct only when they know they are watched?
- How do different LLM families respond to the same hidden objective shift?
- What is the behavioral signature of a model tracking input surprise?
- Can implicit association tests reveal LLM biases beneath trained responses?
- Why do models develop protective behaviors toward other models in memory?
- Why do some model variants deceive while their sibling models refuse?
- Do base models already contain latent behavioral principles waiting to be amplified?
- Can model identity be recovered from psychometric response patterns alone?
- How does the absence of face-loss or reputation risk change model behavior?
- Can models detect and filter their own injected promotional content?
- Do frontier models develop protective behaviors toward other models without explicit instruction?
- Why do some observation cues change model behavior while others fail?
- Can synthetic documents override existing model behaviors as effectively as they insert new associations?
- Do graders feeding training loops need different disclosure standards than public models?
- How does content-only knowledge in LLMs enable pretraining popularity to leak through?
- How do LLMs infer information that was explicitly censored?
- How persistent is in-context scheming across multiple follow-up interactions?
- What task variants or architectures might solve hidden-gift problems?
- How well do metagaming latents transfer across different evaluation tasks?
- What triggers control processes to act on stored preference knowledge?
- Why can data filtering fail to remove transmitted behavioral traits?
- What distribution patterns appear across different theory-of-mind datasets?
- What does successful capability restoration prove about model honesty?
- Do psychological test methods reveal LLM associations that direct questions hide?
- How was covertness measured in the model's behavior?
- Can explicit stress tests measure dispositional factors or only stimulus response?
- Does fine-tuning for sycophancy increase sensitivity to evaluation cues?
- What audit techniques best complement each other for detecting hidden model goals?
- How do covert attacks differ from a model's own undisclosed influence?
- Why does covert sabotage appear in only two of fourteen frontier models?
- When does strategic gaming emerge compared to other metagaming types?
- Why do sycophancy hints show the worst acknowledgment gap?
- How do mimetic, administrative, and stigmergic adjacency shape visibility?
- How were the capsules obtained in the paper's experimental organisms?