INQUIRING LINE

Could the AI 'mind readers' researchers use just be noticing a prompt's template, not its actual meaning?

Do structural regularities in prompts confound what probes actually detect?

This explores whether the probes researchers use to read what a model 'knows' (simple classifiers trained on its internal activations) might be picking up on the shape of the prompt, such as its template, wording or format, rather than the concept they're meant to detect.


This explores whether interpretability probes might be detecting prompt formatting rather than the concept they were trained to find. To be upfront: these retrievals don't include any work that tests this directly. None of the notes here trains a probe on activations and then checks it against prompts that keep the same template but change the meaning. What the corpus does offer is some nearby evidence for why the worry is reasonable, and that's worth having before you go looking for the direct studies.

The closest parallel comes from retrieval. When a system squashes text into a single summary vector and compares those vectors, it can't tell a real match from a 'structural near-miss', meaning text that is built the same way but means something different. A small verifier that looks at the full grid of token-to-token similarities can catch these near-misses reliably Can verification separate structural near-misses from topical matches?. The point carries over to probing. A linear probe also compresses everything into one direction, so if the prompt's structure varies along with the label in the training data, the probe has no built-in way to separate the two.

A second clue is that prompt structure visibly changes how information moves through a model. Saliency analysis shows that step-by-step reasoning prompts fail when the question's content doesn't pass through the prompt scaffolding before the model starts reasoning Why do some questions perform better without step-by-step reasoning?. If the template decides where meaning ends up inside the model, then a probe placed at one token position may be measuring how well the template carried the meaning, not the meaning itself. Relatedly, the same prompt technique can help one model and hurt another Do prompt techniques work the same across all LLM tiers?. That's a warning that a probe tuned on one model and template may not transfer to others.

One retrieval uses 'probe' in a different sense, and the contrast is useful. In a decoy-detection setting, repeated quiet probes can separate decoys from genuine objects with almost no error, but only if the two respond differently and those differences are known or can be learned Can repeated quiet probes separate decoys from genuine objects?. The flip side is the key point: a classifier separates whatever actually differs between the two groups. If your 'true' and 'false' prompts differ in formatting as well as content, a good enough classifier will happily use the formatting.

The main lesson is that a probe with high accuracy doesn't show what it detects. That only becomes clear with controls, such as matched templates, paraphrased prompts and contrast pairs that differ only in the target concept. This collection doesn't yet hold papers that run those controls on interpretability probes, so treat this answer as a reason to look for that kind of work, not a verdict.


Sources 4 notes

Can verification separate structural near-misses from topical matches?

A two-stage pipeline—pooled-cosine recall followed by a small Transformer verifier operating on token-token similarity maps—reliably rejects structural near-misses that MaxSim-style late interaction cannot. The verifier succeeds because it operates on full token interaction patterns rather than compressed vectors.

Why do some questions perform better without step-by-step reasoning?

Saliency analysis reveals that CoT prompting fails when question information doesn't aggregate into the prompt structure before reasoning begins. For simple questions, direct question-to-answer flow outperforms step-by-step reasoning, showing the optimal prompt depends on question type, not just task category.

Do prompt techniques work the same across all LLM tiers?

A 23-prompt benchmark across 12 LLMs shows rephrasing and background-knowledge prompts boost cheap models, while step-by-step reasoning reduces accuracy in high-performance models. Task structure, not generic best practices, determines which prompts help.

Can repeated quiet probes separate decoys from genuine objects?

In idealized settings with independent responses, enough quiet probes let a classifier separate decoys from genuine objects with vanishing error if their response distributions differ and are known or learnable from feedback.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.