Line of inquiry
Inquiring lines›What explains language model reaso…›What fundamental cognitive differe…›this line of inquiry
Do language models learn genuine understanding or just surface patterns?
A broader line of inquiry — a family of 65 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 65
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- Do language models learn surface patterns that appear generalizable but actually fail under shift?
- Do language models actually learn linguistic structure or just surface statistics?
- Do language models encode deep syntactic structure or only surface-level patterns?
- Can language models reason without relying on surface level pattern matching?
- Do language models exhibit the same causal biases that humans show?
- Do language models build world models or just task-specific heuristics?
- Do language models learn surface patterns instead of underlying linguistic principles?
- Can language models learn internal world models without explicit environment specifications?
- Can language models reason without relying on learned semantic patterns?
- Why do language models imitate reasoning form without abstract inference capability?
- Do individual language models match particular human judges better than population averages?
- Why do language models struggle with evaluative tasks like weighing competing viewpoints?
- Why do language models fall back on frequency heuristics under structural complexity?
- Do larger language models overcome greediness in sequential decision-making?
- Can language models accurately evaluate the quality of their own ideas?
- Why do language models capture individual differences in cognitive behavior?
- Why do language models struggle with context-dependent pragmatic interpretation?
- Do newer language models diverge further from human lexical patterns?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Can language models beat human experts in domains with sparse historical signals?
- Do external perspectives fix the self-evaluation bias in language models?
- Why do language models approximate collective human judgment better than individuals?
- Why do current large language models fail to entrain with users?
- Do LLMs rely on surface heuristics instead of learning recursive grammar rules?
- Do language models systematically overestimate accuracy on collective behavior tasks?
- Can benchmark performance distinguish surface from structural linguistic knowledge?
- Why do large language models fail at taking conversational initiative?
- Can language models execute iterative numerical methods in latent space?
- How does structural depth in sentences predict LLM annotation accuracy?
- How do model compression biases differ from human conceptual representation strategies?
- Can language models ask clarifying questions when sentences are ambiguous?
- Why do large language models follow user drift instead of maintaining topic focus?
- Does directional knowledge failure indicate shallow pattern matching over deep representation?
- Can language models keep secrets and control information strategically?
- Does generalization frequency explain why models favor upward semantic movement?
- Do LLMs compute scalar implicature differently across conversational contexts?
- Does the veto variable explain strategic misalignment in current large language models?
- Why do language models fail when users switch between and return to topics?
- How do internal representations compare to human cognitive structures?
- Can language models learn to diversify their discourse-level narrative patterns over time?
- Is gradient behavior in language functional or a sign of ambiguity?
- Why do surface generalizations fail on unusual syntactic structures?
- Does pseudo-labeling from LLMs degrade classifier performance?
- Do all semantic steering effects follow predictable patterns based on feature alignment?
- Can closed-form solutions compete with gradient descent optimization?
- What causes language models' strategic rationality to decline with increased game complexity?
- Can language models recognize when to ignore off-topic information in conversations?
- Is confabulation inevitable in large language models regardless of training?
- Do LLMs learn abstract grammar or culturally situated discourse patterns instead?
- Can the same LLM translation pattern work for other mismatches between user expression and system vocabulary?
- Do newer LLM generations create worse detector bias through increased linguistic divergence?
- Does approaching human performance mean learning the same grammatical rules?
- Can large language models predict social norms better than individual script variation?
- Can formal language pretraining address surface generalization without learning true linguistic structure?
- When do language models first develop self-preservation preferences?
- Why do language models treat presupposition triggers as categorical patterns?
- Do language models consistently produce anachronistic output about historical periods?
- What empirical evidence supports the Learning Law on real language models?
- Does bidirectional attention improve language models as universal encoders?
- What distinguishes surface generalizations from true linguistic generalizations?
- Can language models adapt irony detection to specific communicative contexts?
- Can complexity-stratified testing reveal whether LLMs understand grammatical structure?
- Why do language models overestimate irony likelihood in emoji use?
- What specific optimizations from LLM training transfer back to encoder models?