Line of inquiry
Inquiring lines›What drives capability improvement…›What drives capability improvement…›this line of inquiry
Can mechanistic interpretability methods reliably reveal what models actually know?
A broader line of inquiry — a family of 77 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 77
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can interpretability tools distinguish genuine reasoning from fabricated reasoning inside models?
- Can mechanistic interpretability findings guide practical interventions in model design?
- How do mechanistic interpretability methods surface what models represent internally?
- Can models hide their reasoning in continuous space rather than natural language?
- Do internal belief probes reveal what models actually know versus report?
- How do mechanistic features compare to natural language for interpretability?
- Can attractor dynamics compete with input-based probing for characterizing model knowledge?
- Do models verbalize their implicit knowledge when that knowledge influences their output?
- Can interventions on model components prove mechanism without explaining encoding?
- How do we distinguish knowledge encoding from knowledge usage in models?
- How misaligned are verbal reports from internal model computation?
- What makes representation engineering better than mechanistic interpretability for detecting hidden objectives?
- Can mechanistic interpretability explain explanation-execution disconnection?
- Can models distinguish between activated knowledge and genuine reasoning?
- Can probing reveal whether encoded facts actually influence model outputs?
- Can activation probes detect reasoning that models omit from text?
- What prevents LLM representations from causally influencing generation outputs?
- Can geometric structure in representations exist without supporting functional mechanisms?
- How much do mechanistic interpretability findings reflect true reasoning architecture?
- What is the difference between changing model outputs versus changing internal representations?
- Can models distinguish between injected thoughts and their own outputs?
- Why do feature visualizations alone fail to establish mechanistic claims?
- When does knowledge activation fail across different model architectures?
- Why do models override signals they clearly perceive internally?
- Why do models maintain accurate beliefs but generate false claims?
- What skills can large models identify and organize about their own abilities?
- What distinguishes conceptual understanding from statistical pattern matching in models?
- How does external validation replace the need for model interpretability?
- What distinguishes a representational feature from a causally inert correlation?
- How do probe-based interventions in activation space compare to mechanistic interpretability approaches?
- Can a world model have rich representations without adequate data coverage?
- How do mechanistic interpretability tools help distinguish truthfulness from honesty?
- Does causal intervention alone explain how neural mechanisms implement representations?
- Can representation analysis methods detect complex features models compute with?
- How does mechanistic interpretability reveal ideological structures in language model weights?
- Does sequence prediction accuracy prove an underlying world model exists?
- Can graph cyclicity and topology predict when reasoning systems achieve breakthrough insights?
- Why do models with less steerability have more abstract ideological features?
- Does information stored in neural networks necessarily influence generation decisions?
- How do world models decompose between representation of facts versus generative mechanisms?
- Can causal steering change verbalization without changing internal representation?
- What distinguishes task-specific heuristics from genuine world models?
- Can time-awareness live in model parameters instead of retrieval?
- Do reading vectors from activation space causally control model behavior?
- Does base model geometry predict which associations persist through intervention?
- What makes knowledge editing different from simply finding where facts are stored?
- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- Does next-state prediction alone build mechanistic world models or just sophisticated interpolation?
- Why must world models be nested rather than flat and uniform?
- When does a model's lack of interpretability become a genuine epistemic problem?
- Why does integrating world models with decision-making systems matter?
- Can models be honest without being truthful about facts?
- How do belief edits differ between surface endorsement and deep integration?
- What role does embedding space geometry play in multi-hop reasoning?
- How does LatentQA differ from predefined concept steering like representation engineering?
- What test distinguishes genuine compositionality from fractured feature presence?
- How does representation engineering compare to mechanistic interpretability for auditing?
- How does mechanistic interpretability complement learning mechanics in explaining deep learning?
- How do mechanistic interpretability and scientific understanding relate to each other?
- How does belief-behavior inconsistency relate to instruction execution splits?
- What makes some concepts more steerable than others in activation space?
- Can a theory be justified if the evidence generating it remains opaque?
- What makes a neural network circuit actually interpretable to humans?
- How does the knowing-doing gap relate to Potemkin understanding?
- What does leveraging internal representations during training actually mean operationally?
- Are detection and identification of injections truly separable in neural circuits?
- Why do foundation models fail at hidden state prediction despite sequence accuracy?
- How do you measure the depth of political representation inside a language model?
- What cognitive structures do realistic belief models need to include?
- What makes some internal circuits more interpretable than others?
- What distinguishes new associations from existing ones at the computational level?
- What happens when you remove core political features from a deep model?
- How should disciplines evaluate theories built on opaque machine learning predictions?
- How do ablation studies reveal function without representational characterization?
- How do functional features differ from representational abstract features?
- How do internal activation patterns reveal whether a model is genuinely jailbroken?
- How should world models represent what one person knows versus another?