INQUIRING LINE

A model can 'know' a fact internally and still never use it when answering — why does knowing and using split apart?

When does internal model knowledge fail to appear in outputs?

This explores the gap between what a language model has stored inside it and what it actually says. It asks when and why knowledge the model 'has' never shows up in its answers.


This explores the gap between what a language model has stored inside it and what it actually says. The corpus suggests the gap is common. Having a fact somewhere in the model and using it to write an answer turn out to be separate steps, and either one can fail. Researchers can often read a fact straight out of a model's internal activity and still find that the fact had no effect on what the model produced Do language models actually use their encoded knowledge?. So 'the model knows X' is two claims: the information is in there, and it shapes the output. Only the second one matters to you as a user.

The most practical finding is about reasoning that depends on things nobody said. Models often have the background knowledge a problem needs but don't bring it forward as a constraint. One example: a model knows a car must be present to get it washed, yet misses that this rules out walking to the car wash. The research calls this a bottleneck in noticing which facts apply, not missing knowledge. Gently emphasizing the relevant detail recovers about 15 points of accuracy Why do language models fail to use knowledge they possess?. In a related study, prompting the model to list its unstated assumptions raised accuracy from 30% to 85% Do language models fail at identifying unstated preconditions?. That's an old puzzle from AI research, the 'frame problem' (how a system decides which background facts matter), showing up in a statistical system. The knowledge was there all along. The model just wasn't asking itself the right question.

The flip side is that models do have some internal sense of what they know. Interpretability tools have found a mechanism that detects whether the model recognizes a person or thing. That signal actively decides whether the model answers, refuses, or hallucinates, and it survives into chat-tuned versions Do models know what they don't know?. So a hallucination is sometimes this internal signal firing wrongly, not a total absence of knowledge. Don't expect the model to report this signal accurately if you ask, though. Most of what models say about their own inner states echoes how humans talk in the training data. Real self-report only happens when there's a direct causal link from the internal state to the words Can language models actually introspect about their own states?.

A deeper version of the problem: two models can get identical scores while being organized completely differently inside. One may be coherent. The other may be 'fractured', with knowledge stored in tangled pieces that fall apart under small changes or new situations Can identical outputs hide broken internal representations? Can models be smart without organized internal structure?. Knowledge that works on a test can fail to transfer when the setting shifts, and standard benchmarks won't warn you. The broader overviews treat this as a recurring theme: what's inside a model and what it does on the outside are only loosely linked What actually happens inside large language models? What really happens inside a language model?.

The surprise is that many of the fixes don't add knowledge at all. Light post-training on a small model mostly teaches it how to lay out its reasoning, not new facts, and that alone matches much larger models Can small models reason well by just learning output format?. Systems that learn when to trust their own memory and when to look something up get big accuracy gains, partly by avoiding retrieval they didn't need When should language models retrieve external knowledge versus use internal knowledge?. Cramming more facts into the weights through fine-tuning has a ceiling and can overwrite what the model already knew Can models store unlimited facts without growing larger?. Often the bottleneck isn't how much a model knows. It's whether the right knowledge gets pulled in at the moment it's needed.


Sources 12 notes

Do language models actually use their encoded knowledge?

Multiple studies confirm that language models can encode facts in their representations while those facts fail to causally affect downstream outputs. Encoding and usage are distinct processes.

Why do language models fail to use knowledge they possess?

Models possess relevant knowledge but fail to activate it without explicit prompting. Adding subtle emphasis recovers 15.3 percentage points accuracy, and forcing enumeration of preconditions recovers 6-9 points, showing the bottleneck is in constraint inference, not storage.

Do language models fail at identifying unstated preconditions?

LLMs struggle not from lacking world knowledge but from failing to bring background conditions forward as relevant constraints. Prompting that forces explicit enumeration of preconditions raises accuracy from 30% to 85%, revealing the frame problem persists in statistical systems.

Do models know what they don't know?

Sparse autoencoders revealed that language models develop causal mechanisms for detecting whether they know facts about entities. These mechanisms actively steer both hallucination and refusal behavior, and persist from base models into finetuned chat versions.

Can language models actually introspect about their own states?

LLM self-reports usually reflect human training distributions rather than actual internal processes. However, when a causal chain connects an internal state to accurate reporting—like inferring low temperature from output consistency—genuine lightweight introspection occurs without requiring consciousness.

Show all 12 sources
Can identical outputs hide broken internal representations?

Networks trained with SGD reproduce outputs perfectly while having radically different internal structure than evolved networks, with weight perturbations revealing fractured, entangled representations that prevent transfer to novel contexts or creative recombination.

Can models be smart without organized internal structure?

Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.

What actually happens inside large language models?

Research shows identical accuracy can mask fundamentally different or corrupted internal representations, and mechanistically interpretable circuits may not causally drive outputs. Internal organization and external performance follow distinct paths.

What really happens inside a language model?

Research into mechanistic interpretability, cognitive models, and training dynamics shows that identical benchmark performance conceals radically different internal structures. Improving one capability (helpfulness, accuracy) reliably degrades others (faithfulness, calibration, diversity).

Can small models reason well by just learning output format?

A 1.5B parameter model with LoRA-only post-training matched larger full-parameter RL models on reasoning tasks, suggesting RL teaches output format organization rather than new factual knowledge. This efficiency indicates reasoning and knowledge storage are separable capabilities.

When should language models retrieve external knowledge versus use internal knowledge?

DeepRAG models each reasoning step as a Markov Decision Process where the model learns when to retrieve versus rely on parametric knowledge. The 21.99% improvement comes from better-targeted retrieval and elimination of noise from unnecessary external knowledge.

Can models store unlimited facts without growing larger?

A formal proof and experiments show in-weight memorization is bounded by model size, while tool-use enables unbounded factual recall through a simple circuit. In-weight finetuning also degrades general capability by overwriting prior knowledge.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.