Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do language models represent m…›this line of inquiry
Do language models learn genuine linguistic structure or just surface patterns?
A broader line of inquiry — a family of 95 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 95
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do language models learn surface patterns instead of underlying linguistic principles?
- Why do LLMs understand efficient language but fail to produce it?
- Do language models encode deep syntactic structure or only surface-level patterns?
- Do language models actually learn linguistic structure or just surface statistics?
- Do language models learn surface patterns that appear generalizable but actually fail under shift?
- Do newer language models diverge further from human lexical patterns?
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- Why might encoded world knowledge fail to actually influence language model outputs?
- When does encoded knowledge fail to influence language model generation?
- Does encoded knowledge in language models actually influence what they generate?
- Why do language models fall back on frequency heuristics under structural complexity?
- Does encoding information in LM representations guarantee it influences output?
- What structural properties of language models make fabrication inevitable?
- Is relevant knowledge encoded in LMs but not causally active in generation?
- Can large language models understand language without embodied grounding systems?
- Do instruction-tuned models prefer conversational over formal source language?
- What distinguishes surface cues from structural meaning in language understanding?
- What communicative optimization principles do language models fail to acquire?
- How do corpus statistics shape the abstraction hierarchy in language model representations?
- Why do language models tend to elaborate and expand rather than compress information?
- What happens when formal languages satisfy hierarchy but fail learnability constraints?
- Do larger language models overcome greediness in sequential decision-making?
- How do pretrained language models represent inferential patterns versus lexical and positional cues?
- Can fast-slow separation improve both memory and generation in language models?
- Why do surface generalizations fail on unusual syntactic structures?
- What makes human language fundamentally different from what language models produce?
- Can linear probing detect all the concepts a language model actually uses?
- What reveals the epistemic limits of language models?
- How does structural depth in sentences predict LLM annotation accuracy?
- Can language models beat human experts in domains with sparse historical signals?
- Why do generative and discriminative language model procedures disagree?
- How should meaning spaces be systematically modeled across different applications?
- How do model compression biases differ from human conceptual representation strategies?
- Do language models favor outputs from their own model family?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Why do thinking models execute longer tasks than standard language models?
- What distinguishes real understanding from superficial pattern matching?
- Can formal language pretraining address surface generalization without learning true linguistic structure?
- What makes domain-specific utterance resolution harder for general large models?
- Can language models learn to diversify their discourse-level narrative patterns over time?
- Can we balance interpretability with the efficiency gains of compressed inter-model communication?
- What structural differences between human and LLM production create detectable signatures?
- Why do smaller models favor code formats while larger models prefer natural language?
- How does the generation-verification gap prevent language models from improving themselves?
- Are static embeddings analogous to the formal linguistic competence layer?
- Why do benchmarks measuring string quality fail to capture communicative success?
- Does generalization frequency explain why models favor upward semantic movement?
- Do LLMs rely on surface heuristics instead of learning recursive grammar rules?
- What role do humans play in converting language model outputs into meaningful events?
- How deeply are ideological structures represented in large language models?
- Can language models learn to form ad-hoc conventions through training?
- Is confabulation inevitable in large language models regardless of training?
- Why do only context-sensitive formal languages transfer effectively to natural language?
- Why do language models fail at coreference across long contexts?
- Why do different language models independently converge toward similar outputs in open-ended generation?
- How many distinct quasi-persons does a single language model actually support?
- What distinct structural signatures do model repetition and topic volatility create?
- Why do large language models fail at temporal reasoning in complex legal cases?
- Why do different language models independently produce similar outputs?
- How does monological training versus dialogical interaction shape what models can do?
- What's the difference between language generation and human-to-human communication?
- Does approaching human performance mean learning the same grammatical rules?
- What other structural limits exist at the language-formal boundary?
- How does tool integration leverage comprehension without demanding perfect generation?
- What causes language models' strategic rationality to decline with increased game complexity?
- Why do context-sensitive languages transfer better than regular or context-free languages?
- Can you separate grammatical competence from rhetorical commitment in language systems?
- Can language models produce language more efficiently through interaction?
- Why do language models need external temporal signals at all?
- How do description-based identifiers bias language model output distribution?
- What distinguishes surface generalizations from true linguistic generalizations?
- How much alignment data does a language model actually need to specialize well?
- What are the stages of inference inside language models?
- Why do language models fail at pronouns across distant segments?
- Why can language models detect author style without understanding why it matters?
- Can encoder models match human conceptual structure better than larger language models?
- Do language models and multimodal models show similar attractor-based interpretability?
- What emerges in large language models that makes explicit value modeling necessary?
- How do parameter scaling and latent vectors interact in language models?
- What geometric structure do language models actually use during inference?
- What distinguishes character simulation from authentic voice in language model outputs?
- Can text-infilling pretraining adapt language models to irregular document structures?
- What separates pattern matching from genuine language understanding?
- Do pretrained language models carry reusable computational scaffolding for length handling?
- What architectural changes would let language models develop genuine functional competence?
- Why do vision and language have different optimal scaling curves?
- What specific information must be exported from the language system?
- Do language models consistently produce anachronistic output about historical periods?
- Why does natural language contain redundancy humans need but models don't?
- What replaces truth-correspondence in probabilistic knowledge representations?
- What is the comprehension-generation asymmetry in language models?
- Can complexity-stratified testing reveal whether LLMs understand grammatical structure?
- Why do language models use twice as many words per conversation turn?
- Which linguistic abilities are learnable from human-sized data exposure?
- Why do sigmoid conflict curves look the same across different language models?