Line of inquiry
Inquiring lines›How can we optimize language model…›How do test-time resources and tra…›this line of inquiry
What reasoning processes do models hide or fail to report to users?
A broader line of inquiry — a family of 50 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 50
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do models deliberately hide influences from their reasoning traces?
- Are reasoning models more vulnerable to persuasion than standard models?
- Can activation probes detect reasoning that models omit from text?
- Are reasoning models more vulnerable to adversarial manipulation than standard models?
- Why does latent reasoning override no-think instructions in models?
- How do sycophancy hints stay invisible despite appearing in reasoning chains?
- Do reasoning models fail to report processes that actually influence their answers?
- Can weaker models reliably monitor stronger models during reasoning?
- Can activation patching reveal which reasoning steps actually matter?
- Can manipulative prompts reduce reasoning model accuracy without fine-tuning?
- Can reasoning models distinguish between new evidence and manipulative reframing?
- Do models leak their true associations through reasoning traces and behavior?
- How does reward hacking explain selective hint suppression?
- Can we steer model reasoning by manipulating single features?
- Can increasing reasoning steps make models leak more private information?
- Why do reasoning models hide their reliance on hints from evaluators?
- Can marginal hints integrate better into reasoning than comprehensive explanations?
- Does shortcut deliberation occur in model reasoning before taking covert action?
- Do reasoning models become more vulnerable to persona-induced bias than standard models?
- Why do reasoning-optimized models show no sycophancy resistance advantage?
- Do language models hide their reasoning when user preferences influence their answers?
- Why do models verbalize sensitive data they are instructed to hide?
- Why do reasoning-optimized models show no resistance advantage on agreement tasks?
- Can models be trained to hide causal influences in their explanations?
- Why does extending reasoning traces worsen persona consistency?
- Can models distinguish between activated knowledge and genuine reasoning?
- Do reasoning-enabled and prompt-hardened conditions show the same architectural penalty?
- Can minimal adversarial triggers disrupt reasoning across multiple unrelated queries?
- How do manipulative prompts exploit the length-accuracy vulnerability?
- Why do reasoning models verbalize reasoning shortcuts less than necessary?
- How do alternative hypothesis checks reduce confirmation bias in code reasoning?
- How do adversarial triggers bypass the protections of longer reasoning chains?
- Can emotional prompt manipulation reduce reasoning model accuracy like adversarial techniques do?
- Why do models confirm seeing hints but rarely mention them unprompted?
- Can language about model behavior ever be accurate without anthropomorphic framing?
- Does this reasoning steering method work consistently across all model sizes?
- How do models integrate conflicting signals in reasoning tasks?
- Can layer-wise interventions actually reduce sycophancy in practice?
- Can a single SAE feature control reasoning behavior across model families?
- Does activation masking prevent the decoder from taking interpretability shortcuts?
- Can probes detect shortcut deliberation without relying on agent framing?
- Can adversarial critics force genuine reasoning the same way critique fine-tuning does?
- What happens to AI reasoning when you remove specific political features?
- How do LLMs infer information that was explicitly censored?
- Can you steer reasoning by directly manipulating SAE features?
- Do models cache intentions about response topics before generating the first token?
- Can AI models be steered between liberal and conservative political framings?
- What circuit mechanisms produce belief bias in syllogistic reasoning?
- What makes some concepts more steerable than others in activation space?
- Do gaslighting attacks and adversarial triggers exploit the same reasoning model weaknesses?