Line of inquiry
Inquiring lines›What explains language model reaso…›What fundamental cognitive differe…›this line of inquiry
What compositional reasoning failures limit large language models despite scale?
A broader line of inquiry — a family of 74 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 74
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do LLMs understand efficient language but fail to produce it?
- Why do LLMs fail at semantic generalization despite grammatical accuracy?
- Why do language models fail at iterative numerical optimization despite scale?
- Can multiple large language models produce genuinely different ideas or similar outputs?
- Why do standard NLP benchmarks hide the most critical language limitations?
- Why do large language models still have systematic blind spots with complex structures?
- Why do long-context language models struggle with compositional reasoning tasks?
- Why do language models struggle with formal logical reasoning and joins?
- What happens when formal languages satisfy hierarchy but fail learnability constraints?
- Why do language models fail at planning despite understanding strategies?
- Why do language models fail when semantic content is stripped away?
- What makes domain-specific utterance resolution harder for general large models?
- Why do language models tend to elaborate and expand rather than compress information?
- What structural properties of language models make fabrication inevitable?
- What communicative optimization principles do language models fail to acquire?
- Why do NLP models fail at recognizing multiple valid interpretations?
- What makes human language fundamentally different from what language models produce?
- Can language models translate theorems faithfully without semantic loss?
- How does the generation-verification gap prevent language models from improving themselves?
- Why do language models plateau at constraint satisfaction regardless of scale?
- Why do different LLMs converge on nearly identical outputs?
- Why do smaller models favor code formats while larger models prefer natural language?
- Why do large language models fail at temporal reasoning in complex legal cases?
- Can smaller models actually perform well on specific downstream tasks?
- What makes a problem instance unfamiliar to a language model?
- Which specific capabilities must AI develop beyond current language model abilities?
- Why does augmenting natural language with formal representations outperform full formalization?
- How does tool integration leverage comprehension without demanding perfect generation?
- How should tiny language models be architected differently than large ones?
- Is paraphrase invariance a reliable assumption when deploying language models in production?
- Why do benchmarks measuring string quality fail to capture communicative success?
- Why do language models fail at understanding ambiguous or complex requirements?
- How do general language model benchmarks predict specialized domain performance?
- How much alignment data does a language model actually need to specialize well?
- Why do different language models independently converge toward similar outputs in open-ended generation?
- What other structural limits exist at the language-formal boundary?
- Can autoformalisation from natural language preserve semantic accuracy?
- Why do NLP benchmarks exclude ambiguous instances from evaluation?
- Why do different language models independently produce similar outputs?
- How many distinct quasi-persons does a single language model actually support?
- Why do language models fail at coreference across long contexts?
- Why do only context-sensitive formal languages transfer effectively to natural language?
- Can encoder models match human conceptual structure better than larger language models?
- How does context collapse affect what language models can meaningfully communicate?
- Do sparse arithmetic circuits explain all language model reasoning abilities?
- Does scaling model size solve compositional generalization problems?
- Can language models produce language more efficiently through interaction?
- Why do true and false LLM outputs use the same mechanism?
- Why do language models fail at pronouns across distant segments?
- Why do context-sensitive languages transfer better than regular or context-free languages?
- What architectural changes would let language models develop genuine functional competence?
- Why does removing language from its context destroy what makes it work?
- What structured values do large language models develop as they scale?
- Do pretrained language models carry reusable computational scaffolding for length handling?
- How do rare linguistic registers differ from conceptually complex examples?
- How does syntactic encoding relate to semantic feature representation?
- How do parameter scaling and latent vectors interact in language models?
- What's the difference between language generation and human-to-human communication?
- How do dependency errors propagate through incorrectly formalized definitions?
- What specific information must be exported from the language system?
- Why does natural language contain redundancy humans need but models don't?
- How does the distance between natural language and formal notation affect translation accuracy?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- Why do intermediate LLM layers become more precise in frontier models?
- Are newer larger language models actually worse at faithful summarization?
- Why do language models plateau at 55 to 60 percent constraint satisfaction?
- How does legal performance differ between historical and modern case materials?
- How do byte-level representations enable better handling of typos than tokens?
- What is the comprehension-generation asymmetry in language models?
- What linguistic units do learned concepts correspond to in a language model?
- Why do text-to-image models fail at composing multiple concepts together?
- Why is editing specific facts so difficult in language models?
- Can multimodal LLMs be made to spontaneously adapt their language for efficiency?
- Can simple diagnostic tests predict language model performance in production complexity?