INQUIRING LINE

When a model ignores what you tell it, the fix often lies inside its internal workings, not in better prompts.

Do uninterpretable learned representations create robustness problems in language models?

This explores whether the fact that we can't easily read what a language model has learned internally makes it brittle: likely to fail in ways nobody can see coming or fix.


This explores whether the opacity of what models learn internally turns into brittleness: failures nobody can predict, diagnose, or fix. One caveat up front: the collection has little that addresses interpretability head-on. What it does have points to an interesting answer. Learned representations do cause robustness problems. But you don't always have to read the internals to anticipate those problems, and when you do want to fix them, the representations themselves are often where you have to work.

The clearest case is a model ignoring what you tell it. When knowledge absorbed during training is strongly held, models produce answers that contradict the information right in front of them. Better prompting can't talk them out of it. Researchers had to intervene directly on the internal representations to make the context win Why do language models ignore information in their context?. That's the robustness cost of opacity in one picture: the problem sits in a layer that ordinary instructions can't reach. A similar pattern shows up in grammar. Models that handle everyday text fluently consistently misidentify embedded clauses and complex phrases, and they get worse as sentence structure gets deeper Why do large language models fail at complex linguistic tasks?. The same goes for ambiguous sentences, where GPT-4 sorts out the competing meanings only about a third as often as humans do Can language models recognize when text is deliberately ambiguous?. Both look like what happens when learned representations capture surface statistics rather than the underlying rules. Ordinary benchmarks hide that gap until someone tests for it.

The surprising counterpoint is that you can sometimes predict failures without opening the black box at all. One line of work treats an LLM simply as a machine trained to predict the next most probable word. From that description alone, researchers correctly predicted that tasks with improbable-looking answers, like reciting the alphabet backwards or counting letters, would trip models up even though they're logically trivial Can we predict where language models will fail?. Confidence works the same way from the outside. Instead of trusting what a model 'feels' internally, one method looks up how the model did in the past on cases where it was similarly confident, and that matches far more expensive checks Can past performance predict when a model will be right?. Uninterpretable internals don't leave us blind. They push us toward reasoning about behavior.

The twist is that not every strange internal change is a failure. When models face unfamiliar, out-of-distribution tasks, their internal activity becomes markedly sparser. You might read that as the model breaking down, but the evidence suggests it's adaptive filtering that helps hold performance steady Do language models sparsify their activations under difficult tasks?. So the robustness problem isn't simply that representations are hard to read. It's that without reading them, we can't tell a coping mechanism from a breakdown. If you want to see where robustness gets designed in rather than discovered afterward, the work on teaching models to abstain when uncertain is a good next stop Can models learn to abstain when uncertain about predictions?.


Sources 7 notes

Why do language models ignore information in their context?

Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.

Why do large language models fail at complex linguistic tasks?

Top-tier LLMs like Llama3-70b consistently misidentify embedded clauses, verb phrases, and complex nominals. Performance degrades predictably as syntactic depth increases, revealing that statistical learning captures surface patterns but not deep grammatical rules.

Can language models recognize when text is deliberately ambiguous?

AMBIENT benchmark shows GPT-4 correctly disambiguates only 32% of cases versus 90% for humans. This failure spans lexical, structural, and scope ambiguity—revealing that LLMs cannot hold multiple interpretations simultaneously, a fundamental gap hidden by standard benchmarks.

Can we predict where language models will fail?

By framing LLMs as autoregressive probability machines, researchers predicted tasks with low-probability target responses would be systematically harder, even when logically simple. Experiments confirmed predictions like backwards alphabet and letter counting.

Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

Show all 7 sources
Do language models sparsify their activations under difficult tasks?

As task difficulty increases, LLM hidden states become substantially sparser in a localized, systematic way that correlates with task unfamiliarity and reasoning load. This sparsification acts as a selective filter stabilizing performance under OOD shift rather than a failure mode.

Can models learn to abstain when uncertain about predictions?

Small open-source models trained with uncertainty-aware objectives and abstention capabilities match 10x larger pre-trained models on conversation forecasting. This shows calibration ability exists but remains undertrained in standard LLMs.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.