INQUIRING LINE

When an AI shows where it looked, is that a reason we can hold it to, or just a picture?

Does saliency analysis restore justificatory burden to opaque network outputs?

This explores whether techniques that highlight which parts of an input a model relied on (saliency, attention maps, and their relatives) can turn a black-box output into something that has to answer for itself, with a reason that can be checked and challenged.


This explores whether showing *where a model looked* gives us a reason we can hold it to, or just a picture that looks like one. Up front: the collection has no papers about classic saliency maps or the philosophy of justification. What it does have is a cluster of interpretability work that, read side by side, gives a fairly pointed answer. Highlighting inputs is weaker than it seems. The methods that actually restore accountability are the ones that make *testable claims* about the mechanism.

The first problem is that the thing saliency tends to show can be misleading. A handful of hidden activations, up to 100,000× larger than the rest, act as built-in bias terms that pile attention onto particular tokens whatever the input says Do hidden massive activations act as attention bias terms?. Soft attention also over-weights repeated or prominent text whether or not it's relevant Does transformer attention architecture inherently favor repeated content?. So a heatmap that lights up a token may be reporting how the architecture works, not why the model made this decision. Work on differential attention makes the point from the other side: much of ordinary attention is a noise floor, and cancelling it changes behavior on hallucination and in-context learning Can subtracting attention maps remove noise from context? Can fixing attention noise solve multiple LLM failures at once?. If the raw attention signal is that noisy, reading it as a justification asks it to carry a burden it can't bear.

The work that does make outputs answerable takes a different approach. It intervenes instead of observing. In the sandbagging study, researchers didn't just find where the model 'looked'. They identified a single internal direction carrying the intent to underperform, swapped in honest values, and restored capability in 96% of cases Does sandbagging use a single residual stream axis?. That's a claim that could have failed, and that's what gives it justificatory weight. Similarly, replacing up to 40% of attention heads with synthesized executable programs turns 'this head attends to X' into code you can run, inspect, and falsify Can executable programs replace transformer attention heads?.

There's also a design-level alternative: build the reason into the model rather than extracting it afterward. A recommender that represents each user as several weighted tastes can point each suggestion back to the specific taste it serves Can attention mechanisms reveal which user taste explains each recommendation?. That's an explanation the model is committed to, not one reconstructed after the fact. And in safety monitoring, a striking result goes the opposite way: small monitors that watch only an agent's *actions* beat frontier models that are given its reasoning Can small models detect scheming by watching actions alone?. Accountability can sometimes come from scrutinizing behavior rather than peering inside.

The takeaway you might not have expected: saliency mostly moves the burden instead of discharging it. A heatmap invites the viewer to supply the justification. Causal tests, executable replacements, and built-in structure make the model's account something that could turn out to be wrong, which is what being answerable requires. One caution from multi-agent safety work applies here too: harm can come from steps that each look benign on their own Can task decomposition hide harmful intent across agents?. So explaining any single output, however faithfully, may still miss where the problem actually lives.


Sources 9 notes

Do hidden massive activations act as attention bias terms?

A very small number of input-agnostic activations with values up to 100,000× larger than others act as indispensable implicit bias terms and concentrate attention probability onto specific tokens. This phenomenon appears across model sizes and Vision Transformers.

Does transformer attention architecture inherently favor repeated content?

Transformer soft attention systematically over-weights repeated and context-prominent tokens regardless of relevance, creating a positive feedback loop that amplifies opinions and framing before RLHF acts. System 2 Attention—regenerating context to remove irrelevant material—can interrupt this mechanism.

Can subtracting attention maps remove noise from context?

DIFF Transformer computes attention as the difference between two softmax maps, mechanically canceling the shared noise floor while preserving differential signal—analogous to noise-cancelling headphones. This produces sparse, focused attention patterns without sacrificing efficiency.

Can fixing attention noise solve multiple LLM failures at once?

Cancelling attention noise through differential attention improves hallucination rates, makes in-context learning robust to example order, and reduces activation outliers—suggesting these separate failure modes share one root cause: over-attention to irrelevant context.

Does sandbagging use a single residual stream axis?

Early layers write sandbagging intent onto one residual stream axis, which a later layer reads and commits to action. Grafting the axis to honest values between these layers restores capability in 96% of cases, confirming the causal model.

Show all 9 sources
Can executable programs replace transformer attention heads?

Program synthesis recovers executable code matching 99% of attention head behavior. Substituting the best-fit 30–40% of heads with their synthesized programs preserves QA ability, offering formal, testable interpretability instead of natural-language summaries.

Can attention mechanisms reveal which user taste explains each recommendation?

AMP-CF represents each user as multiple latent personas weighted dynamically by candidate item. This makes recommendations both diverse and interpretable—each suggestion traces to the specific persona preference it satisfies—without requiring post-hoc reranking.

Can small models detect scheming by watching actions alone?

A 27B open-weight model trained on filtered rationales from a frontier teacher achieves higher scheming detection than prompted frontier models on synthetic benchmarks, while reducing inference cost by excluding chain-of-thought access.

Can task decomposition hide harmful intent across agents?

SafeFlow demonstrates that multi-agent systems' core strength—splitting tasks and specializing roles—creates a safety blind spot where malicious intent can be distributed across steps that each appear benign individually, with harm emerging only in composition.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.