Line of inquiry
Inquiring lines›What explains language model reaso…›Where and why do LLM capabilities…›this line of inquiry
Why don't LLMs reliably translate capability into accurate outputs?
A broader line of inquiry — a family of 99 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 99
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Why does LLM knowledge fail to influence their actual outputs?
- How faithful are natural language explanations from LLMs really?
- Do standard language benchmarks underestimate what LLMs can actually do?
- Why do LLMs generate novel ideas but struggle to evaluate them?
- Can LLMs reliably assess the quality of ideas they generate?
- Why do LLMs excel at generation but struggle with evaluation?
- Do LLMs struggle more with semantic accuracy than syntactic correctness across domains?
- Why do language models produce plausible outputs over accurate failure reports?
- Can surface-level correctness hide failures in structural learning by LLMs?
- What makes novelty assessment harder to automate than idea generation?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- Why do NLP benchmarks hide LLM failures in ambiguity handling?
- Do LLMs fail exploration because of context integration or computational limitations?
- Can LLMs reliably audit other language models for errors?
- What structural differences between human and LLM production create detectable signatures?
- What happens when we treat LLM outputs as sampled rather than stored?
- Should LLMs query users back when presented with under-specified scenarios?
- Why do benchmark tests fail to detect LLM comprehension gaps?
- Can users experience the LLM Fallacy even when AI outputs are completely accurate?
- Why do LLMs choose incorrect edits despite understanding the task?
- How do constrained versus unconstrained domains flip LLM novelty patterns?
- How can a model explain something correctly yet fail to apply it?
- Why do backward-looking benchmarks underestimate LLM scientific value?
- Can critique-only calls in LLMs exploit a measurable gap between generation and evaluation?
- Can LLMs generate more novel research ideas than human experts?
- What structural barriers prevent LLMs from making evaluative judgments about writing?
- Can lightweight verification methods help experts trust LLM outputs?
- Why do LLMs struggle more when only numerical values change?
- Where do LLMs fail as knowledge systems compared to humans?
- Why do LLMs fail at directly solving stochastic control problems?
- Can LLMs recognize rhetorical devices they cannot actually produce themselves?
- Do language-model agents reach more accurate conclusions on objective versus subjective questions?
- Where do LLMs succeed at generation but struggle with evaluation?
- Does prompting for accuracy actually reduce LLM hallucinations and errors?
- What capability boundary exists in LLM prediction of effect sizes?
- Why do LLMs generate ideas that sound novel but fail during execution?
- Can LLMs explain concepts correctly while failing to use them?
- Why do users systematically overrely on confident LLM outputs across languages?
- Can researchers prevent their expectations from shaping LLM outputs?
- Why do LLMs explain evidence accurately while missing its implications?
- Which knowledge types do LLMs handle better than humans in reasoning tasks?
- Why do LLMs fall for and deploy logical fallacies with equal confidence?
- Can language models reliably score open-ended collaboration discussions against skill rubrics?
- Why does analytical depth demand trigger fabrication over transparent uncertainty?
- Do LLMs generate more novel ideas than they can evaluate?
- Why do people misinterpret or misuse LLM outputs in practice?
- What specific execution barriers do LLM ideas encounter most frequently?
- Why do models generate creative ideas but fail to evaluate their legitimacy?
- How can LLMs evaluate their own creative outputs for utility and novelty?
- Can auditing LLM performance on complex inputs improve NLP pipeline reliability?
- Do LLMs detect harmful concepts before they influence model outputs?
- Does exposure to more domain-specific examples reduce LLM overconfidence?
- Why do LLMs plateau on creativity tasks while humans reach further?
- What levels of understanding about LLM knowledge representation can automated systems reliably extract?
- What distinguishes planning knowledge from an executable plan that works?
- Why do LLMs fail at iterative numerical computation in latent space?
- Does LLM miscalibration cause failures in clinical information extraction?
- Can external summarization solve exploration problems in complex real-world environments?
- Why do LLMs generate novel ideas but lack evaluative commitment?
- Can an LLM be well calibrated but still unreliable on single evaluations?
- How do human feedback and data distribution shape LLM discourse competence?
- Why does single-turn Q&A framing not match real user deployment patterns?
- Why do LLM-generated ideas score higher novelty yet lower feasibility than expert ideas?
- Which use cases can tolerate unverified LLM outputs without external verification?
- Why do LLM outputs need verification even when they look polished?
- Why does probability of text completion not equal knowledge value?
- Can we systematically enumerate LLM failure modes from first principles?
- How does removing a spurious cue change LLM performance?
- What workflow structure pairs LLM generation with human evaluation most effectively?
- Can structured decomposition fix evaluation gaps in other research tasks?
- Can tool use or self-conditioning fix long-horizon delegation drift in LLMs?
- Why do leaderboard metrics fail to capture human flourishing in LLM evaluation?
- How can we verify outputs from systems that generate without grounding?
- Do longer prediction horizons systematically degrade LLM forecasting accuracy?
- What distinguishes entity errors from relation errors in LLM output?
- How do LLMs compress specific expert knowledge into median abstraction?
- What causes LLMs to ignore unstated constraints they know about?
- Why does direct generation work better for some deliverable types than others?
- Can prompted or fine-tuned models generate genuine narrative ambiguity?
- What makes task alignment more fragile than underlying knowledge retention?
- Which LLM backends produce the most executable research ideas?
- Why does regenerating LLM responses produce different but equally valid answers?
- How does the inability to manage ambiguity undermine literary analysis tasks?
- Why do LLM explanations feel authoritative even when alignment with the model fails?
- Why do experts experiencing the LLM Fallacy fail to develop custodian skills?
- What barriers prevent experts from specifying concepts for LLM extraction?
- How long does retrievability support error detection across repeated LLM use?
- Can output-layer corrections fix fundamental cultural representation deficits in LLMs?
- Why do LLMs produce directive responses when experts favor open-ended exploration?
- How can analysts customize generated UIs without learning to think like engineers?
- Why do LLM personas struggle with specificity in specialized domains like law?
- Why does LLM fluency create false perceptions of professional standing and expertise?
- Do LLMs match top human creative writers in literary quality?
- How does output homogeneity across different LLMs compare to narrowness within a single model?
- What role do model-based critics play in validating LLM plans?
- What explains the 87 percent to 12 percent cliff in plan executability?
- What causes silent document corruption in long LLM workflows?
- What happens when experts prompt using their own technical register?