Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do modularity, routing, and se…›this line of inquiry
What critical LLM failures do standard benchmarks hide?
A broader line of inquiry — a family of 34 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 34
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do standard NLP benchmarks hide the most critical language limitations?
- Do standard language benchmarks underestimate what LLMs can actually do?
- Why do language models fail at iterative numerical optimization despite scale?
- Why do NLP benchmarks hide LLM failures in ambiguity handling?
- Why do language models plateau at constraint satisfaction regardless of scale?
- Why do large language models still have systematic blind spots with complex structures?
- What prevents monolithic LLMs from coordinating decomposition with execution?
- Do LLMs fail exploration because of context integration or computational limitations?
- Why do benchmark tests fail to detect LLM comprehension gaps?
- Why do LLMs fail at iterative numerical computation in latent space?
- Why do language models fail at planning despite understanding strategies?
- Why do NLP benchmarks systematically exclude ambiguous test cases from evaluation?
- Why do NLP benchmarks exclude ambiguous instances from evaluation?
- Can language models execute iterative numerical methods in latent space?
- Why do language models fail at understanding ambiguous or complex requirements?
- How do general language model benchmarks predict specialized domain performance?
- Why do naive pruning and quantization destroy LLM performance so easily?
- Why do LLMs fail at directly solving stochastic control problems?
- Why do LLMs struggle more when only numerical values change?
- Why do LLMs degrade on long inputs before hitting context limits?
- Why do different LLMs converge on nearly identical outputs?
- Does the alignment frame mislead us about what LLM problems actually are?
- Why do language models plateau at 55 to 60 percent constraint satisfaction?
- Can auditing LLM performance on complex inputs improve NLP pipeline reliability?
- How do LLM activations sparsify differently under out-of-distribution inputs?
- Why do rare complex structures in training data harm LLM generalization?
- What constraint satisfaction rate do LLMs achieve at scale?
- What mechanism causes LLMs to plateau on numerical optimization tasks?
- What latent mechanisms do LLMs use when they cannot execute iterative methods?
- Why does fixing decomposition step count matter more than vocabulary alignment?
- Why do intermediate LLM layers become more precise in frontier models?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- Are newer larger language models actually worse at faithful summarization?
- Why does genetic programming outperform direct LLM generation by 86 percent?