Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
How do training priors constrain what context information can override?
A broader line of inquiry — a family of 56 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 56
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does context information fail to override prior training associations?
- Does foundational model training or user priors more strongly shape final outputs?
- How do training associations override context information in language models?
- How does training order affect knowledge acquisition in language models?
- Can models recover knowledge with completely unrelated retraining tasks?
- How do training-data priors influence model defaults when context is ambiguous?
- Can in-context learning substitute for domain-specific training altogether?
- Can data filtering during pretraining prevent cognitive biases in language models?
- Can knowledge encoded in model representations fail to influence generation?
- Why do language models substitute parametric knowledge over retrieved context mid-reasoning?
- Does attention bias explain grounding failure in language models?
- Can prompt-based debiasing work if biases are embedded in pretraining?
- Do instruction-tuned models learn tasks or just output format distributions?
- Do negative constraints require fundamentally different training signals than positive instructions?
- Why is in-context learning brittle to the order of examples presented?
- Why do pretrained model priors reduce the usefulness of retrieved experience?
- Why does consistency training make models resistant to prompt perturbations?
- Why does training data saliency distort how models judge meaning?
- How does parametric knowledge sabotage context-grounded question answering?
- How do training data distributions constrain what language models can accurately know?
- Can in-context learning's advantage erode once interaction histories exceed the context window?
- How do model priors enable targeted context queries without full attention?
- Why does evaluating errors teach more than imitating correct responses?
- What makes some contexts learnable as rules versus requiring model retraining?
- Can explicit numerical signals override learned linguistic defaults in fine-tuned models?
- Why does negative experience transfer better than positive examples alone?
- Can models be trained to hide causal influences in their explanations?
- Can retrieval policies learn to use pretraining statistics as decision features?
- How does in-context learning trigger phase transitions in model behavior?
- Can goal information injected at inference time replace goal-conditioned training?
- Should user context live in tokens or in learned model representations?
- Can Q-priming further strengthen clarifying question behavior beyond social meta-learning alone?
- Does training on critiques of noisy responses produce deeper understanding than imitating correct ones?
- Can models internalize retrieved context as static parametric knowledge?
- Can priming from different facts interfere with each other in the same model?
- How do label constraints improve synthetic data without ground truth validation?
- How much can mitigation techniques like augmentation reduce priming without harming learning?
- What is the difference between changing model outputs versus changing internal representations?
- How does keyword priming enable language models to spread poisoned information?
- Why do structure-targeted training negatives fail to fix the underlying problem?
- Can models converge on similar experience descriptions across different architectures?
- Can neural networks learn that A implies B in reverse?
- How does training distribution shape what language models understand best?
- What causes catastrophic forgetting during domain knowledge embedding?
- How should training data be constructed to preserve teacher-student information gaps?
- Why does monological training prevent models from overriding statistical priors?
- Why do Generation-Then-Comprehension and AI Delegation produce opposite learning outcomes?
- Why does teacher forcing fail to capture long-range dependencies?
- Can a rejected-edit buffer work like hard negatives in contrastive learning?
- How do LLMs infer information that was explicitly censored?
- How would you redesign context integration to prevent prior associations from dominating?
- How do early layers preserve unbiased information while late layers conform?
- What explains the contextual variability of knowledge in transformers?
- Can implicit linguistic information ever be reliably learned from training data?
- Why does keyword priming require only three training exposures to establish?
- What mechanism makes keyword probability the strongest predictor of priming?