Line of inquiry
Inquiring lines›How can we optimize language model…›How do test-time resources and tra…›this line of inquiry
How do language models integrate parametric and contextual knowledge?
A broader line of inquiry — a family of 36 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 36
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do language models substitute parametric knowledge over retrieved context mid-reasoning?
- Can models internalize retrieved context as static parametric knowledge?
- Can models recover knowledge with completely unrelated retraining tasks?
- Does finetuning facts into weights overwrite existing model capabilities?
- How does parametric knowledge sabotage context-grounded question answering?
- What makes some contexts learnable as rules versus requiring model retraining?
- How do we distinguish knowledge encoding from knowledge usage in models?
- What mechanism transfers explicit memories into parametric model weights?
- Why do pretrained model priors reduce the usefulness of retrieved experience?
- What causes catastrophic forgetting during domain knowledge embedding?
- How do training-time and inference-time knowledge injection techniques compare?
- How do layer-wise versus parameter-wise merging strategies affect information retention?
- How does in-weights adaptation create spurious forgetting in models?
- When does knowledge activation fail across different model architectures?
- How do trained weights differ from a stored library or text?
- How do newly learned facts become accessible after gradient updates?
- Should user context live in tokens or in learned model representations?
- Is forgetting in language models reversible or permanent knowledge loss?
- How do language models treat injected evidence as shared background knowledge?
- How should rapidly evolving domains choose knowledge injection methods?
- What training cost tradeoffs exist between fine-tuning and other knowledge injection methods?
- Can time-awareness live in model parameters instead of retrieval?
- What makes knowledge editing different from simply finding where facts are stored?
- Can priming from different facts interfere with each other in the same model?
- Can synthetic documents override existing model behaviors as effectively as they insert new associations?
- What replaces truth-correspondence in probabilistic knowledge representations?
- What causes overfitting when forcing new facts into model weights?
- What explains the contextual variability of knowledge in transformers?
- How would you redesign context integration to prevent prior associations from dominating?
- Which domains need knowledge injection versus reasoning-focused training?
- Can context windows and RAG actually change what language models generate?
- What techniques work best for injecting domain knowledge at training time?
- What non-parametric methods could replace latent factors for inductive learning?
- Why is editing specific facts so difficult in language models?
- What role does knowledge injection play in adapting RAG to industry taxonomies?
- How does upward distillation transfer knowledge from smaller to larger networks?