Line of inquiry
Inquiring lines›Why are language models fragile de…›How do learned model representatio…›this line of inquiry
Do language models learn genuine linguistic structure or just surface statistical patterns?
A broader line of inquiry — a family of 72 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 72
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do language models learn surface patterns instead of underlying linguistic principles?
- Why does context information fail to override prior training associations?
- Do language models learn surface patterns that appear generalizable but actually fail under shift?
- Do newer language models diverge further from human lexical patterns?
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- Do instruction-tuned models learn tasks or just output format distributions?
- Do language models actually learn linguistic structure or just surface statistics?
- Does foundational model training or user priors more strongly shape final outputs?
- Can benchmark performance distinguish surface from structural linguistic knowledge?
- Why might encoded world knowledge fail to actually influence language model outputs?
- Does encoded knowledge in language models actually influence what they generate?
- Do language models encode deep syntactic structure or only surface-level patterns?
- How does training order affect knowledge acquisition in language models?
- Do instruction-tuned models prefer conversational over formal source language?
- When does encoded knowledge fail to influence language model generation?
- Does attention bias explain grounding failure in language models?
- Can knowledge encoded in model representations fail to influence generation?
- Can formal language pretraining address surface generalization without learning true linguistic structure?
- Can large language models understand language without embodied grounding systems?
- How do pretrained language models represent inferential patterns versus lexical and positional cues?
- Why do surface generalizations fail on unusual syntactic structures?
- Is relevant knowledge encoded in LMs but not causally active in generation?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Can language models acquire meaning from distributional patterns alone without joint attention?
- How do training associations override context information in language models?
- How do language models transmit traits through semantically unrelated data?
- Why do only context-sensitive formal languages transfer effectively to natural language?
- Is gradient behavior in language functional or a sign of ambiguity?
- What communicative optimization principles do language models fail to acquire?
- How can language models extract more value from fewer demonstrations?
- Can in-context learning substitute for domain-specific training altogether?
- Does the prediction unit shape what language models actually learn?
- What happens when formal languages satisfy hierarchy but fail learnability constraints?
- Why does training data saliency distort how models judge meaning?
- Can interventions on individual features reliably steer language model behavior?
- Do representations in models causally influence text generation?
- Does next-token prediction alone produce genuine functional language competence?
- Does approaching human performance mean learning the same grammatical rules?
- Can explicit numerical signals override learned linguistic defaults in fine-tuned models?
- Are static embeddings analogous to the formal linguistic competence layer?
- Does generalization frequency explain why models favor upward semantic movement?
- Can structural perturbations harm model accuracy more than semantic ones?
- Do base models or search agents win at open-ended prediction tasks?
- Do negative constraints require fundamentally different training signals than positive instructions?
- Why do context-sensitive languages transfer better than regular or context-free languages?
- Can language models learn to form ad-hoc conventions through training?
- What does next-token prediction tell us about compositional linguistic competence?
- How do humans learn language through communication differently than LLM text prediction?
- Can training on text corpora teach what communicative acts produce?
- Can we reverse the instruction-following deficit through targeted training?
- Can in-context learning's advantage erode once interaction histories exceed the context window?
- Can goal information injected at inference time replace goal-conditioned training?
- What distinguishes surface generalizations from true linguistic generalizations?
- Can instance seeds work for tasks beyond language understanding benchmarks?
- What empirical evidence supports the Learning Law on real language models?
- Can presupposition projection strength vary by context in embeddings?
- How does in-context learning trigger phase transitions in model behavior?
- Why do language models need external temporal signals at all?
- Why do language models treat presupposition triggers as categorical patterns?
- Can statistical learning from language alone capture all aspects of cultural competence?
- How does keyword priming enable language models to spread poisoned information?
- Can neural networks learn that A implies B in reverse?
- Can text-infilling pretraining adapt language models to irregular document structures?
- Can models converge on similar experience descriptions across different architectures?
- Can language models produce language more efficiently through interaction?
- Can statistical learning from text replace embodied cultural experience?
- Can implicit linguistic information ever be reliably learned from training data?
- Do pretrained language models carry reusable computational scaffolding for length handling?
- Why does hierarchical formal language training improve token efficiency more than natural language?
- Why does teacher forcing fail to capture long-range dependencies?
- What causes gradient-based steering via natural language descriptions to work?
- Which linguistic abilities are learnable from human-sized data exposure?