INQUIRING LINE

Does an AI have to be understood from the inside before it can teach us anything about the world?

How do mechanistic interpretability and scientific understanding relate to each other?

This explores two connected questions: whether mechanistic interpretability, the work of reverse-engineering what happens inside a model, is itself a kind of science, and whether we need to understand a model's insides before we can use it to understand the world.


This explores whether mechanistic interpretability is a science in its own right, and whether opening up a model is required before it can teach us anything about the world. The corpus answers the two halves differently. Interpretability is starting to look and work like an empirical science. But using AI to produce scientific knowledge may not depend on interpretability as much as people assume.

On the first half, the clearest point is that interpretability borrows the standards of experimental science. One method alone is not enough. Finding where a concept shows up in a model's internal activity gives you a correlation. Intervening to change the model's behavior gives you an effect, but not an explanation. Only the pairing counts as understanding: first locate a candidate, then test it causally (Can LLM understanding rely on just representation or causation alone?). That is the same move neuroscience and psychology have made for decades. One line of work argues that cognitive science's toolkit transfers almost directly. Its core is Marr's three levels: what problem is being solved, what procedure solves it, and what hardware runs it. These let you explain a model in layers instead of hunting for one master explanation (Can cognitive science methods unlock how LLMs actually work?). The science is also being automated. An agentic system called Mechanist draws on a knowledge graph of about 13,000 prior studies and a library of 32 methods. It proposes hypotheses about mechanisms and runs the experiments to test them, beating general AI-scientist baselines (Can AI automate the discovery of how AI models work?). So AI is now doing science on AI.

Interpretability also feeds a philosophical question: what would it mean for a model to understand something? One account uses interpretability findings to describe three tiers. The first is concepts, which show up as directions in the model's internal space. The second is facts about the world, which show up as connections between those concepts. The third is principles, which show up as compact circuits. The higher tiers sit alongside cheap shortcuts instead of replacing them, so the result is a patchwork (Do language models understand in fundamentally different ways?). That patchwork matters for science. A model can apply a real principle on one input and a shallow pattern on the next, and its outputs alone won't tell you which happened. Measures like the deep-thinking ratio try to read this from the inside. It tracks how often the model's predictions get revised as they pass through its layers (Can we measure how deeply a model actually reasons?).

The second half of the question brings the surprise: science may not need interpretability at all. One argument holds that a model's opacity only matters when you treat its output as a justified claim. If the model just points you toward a hypothesis, the justification comes from the theory you build and test afterward, by your field's normal standards (Can opaque models guide discovery without needing interpretation?). Neuroscience offers a striking case. Fine-tuned LLMs beat human experts at predicting which experimental results actually occurred. The same habit of blending patterns that causes hallucination in look-up tasks turns into useful prediction when the task looks forward (Can LLMs predict novel scientific results better than experts?). A three-role framework for AI in science makes the gap explicit. AI already works as a computational tool and as a source of ideas, but nothing yet shows it understanding science as an independent agent (Can artificial intelligence ever truly understand science?).

There is a twist. Interpretability is drifting toward natural language: models explaining models, sometimes auditing each other (Can natural language explanations redefine what interpretability means?). This greatly expands how much can be explained, but it brings a familiar risk: explanations that sound right without reflecting what the model actually did. Another note argues that whether any explanation works depends on who presents it, how it is framed, and who receives it, and not only on its content (What if XAI is fundamentally a communication problem?). So human understanding is still needed at the end of the chain. Interpretability can show what is inside a model, but someone still has to make sense of it.


Sources 10 notes

Can LLM understanding rely on just representation or causation alone?

Research shows that representational analysis alone identifies correlates without proving causation, while causal analysis alone demonstrates effects without explaining function. Only paired methodology—locating candidates representationally then verifying causally—produces genuine mechanistic understanding rather than descriptive claims.

Can cognitive science methods unlock how LLMs actually work?

Cognitive science's 70-year toolkit of behavioral probes, causal interventions, and representational analysis transfers directly to LLM interpretation. Marr's computational, algorithmic, and implementation levels reframe the problem structurally and enable layered rather than monolithic explanation.

Can AI automate the discovery of how AI models work?

Mechanist, an agentic system pairing a 13,000-study knowledge graph with 32 foundational methods, generates higher-quality mechanism hypotheses and executes experiments more reliably than existing AI-scientist baselines. Four case studies demonstrate discovery of new model behaviors and mechanism-guided interventions.

Do language models understand in fundamentally different ways?

Mechanistic interpretability reveals conceptual understanding (features as directions), state-of-world understanding (factual connections), and principled understanding (compact circuits). Crucially, higher tiers coexist with lower-tier heuristics rather than replacing them, creating a patchwork of capabilities.

Can we measure how deeply a model actually reasons?

Deep-thinking ratio (DTR) measures the proportion of tokens whose predictions undergo significant revision across model layers, correlating robustly with accuracy across AIME, HMMT, and GPQA benchmarks. Think@n, a test-time strategy using DTR, matches self-consistency performance while reducing inference costs.

Show all 10 sources
Can opaque models guide discovery without needing interpretation?

Deep learning models can guide discovery through opaque outputs without interpretation because justification applies to the resulting theory, not the model. Two cases show accurate predictions leading to theories that pass disciplinary standards independent of model understanding.

Can LLMs predict novel scientific results better than experts?

BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.

Can artificial intelligence ever truly understand science?

A framework distinguishes three roles for AI in science: as a computational tool, a source of ideas, and as an independent agent. Evidence shows AI succeeds in the first two but has not yet achieved understanding in the third role.

Can natural language explanations redefine what interpretability means?

LLMs' capacity to explain in natural language expands the scale and complexity of patterns conveyable to humans, enabling ambitious new interpretability goals including model-to-model auditing. However, this medium introduces critical risks: hallucinated explanations that feel plausible but lack faithfulness.

What if XAI is fundamentally a communication problem?

Explanation quality is not intrinsic to the explanation itself but depends on the rhetorical situation: who presents it, how it is framed, and what role the recipient plays. Evaluations that ignore this triad measure only a narrow slice of real-world effectiveness.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.