Line of inquiry
Inquiring lines›Why are language models fragile de…›Why do linguistic mismatches cause…›this line of inquiry
Do LLM explanations accurately predict LLM outputs?
A broader line of inquiry — a family of 80 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 80
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How faithful are natural language explanations from LLMs really?
- Does LLM reasoning always match the outputs it generates?
- Why does LLM knowledge fail to influence their actual outputs?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Can evidence density alone shift an LLM from generation to reasoning?
- When should an LLM engage extended reasoning versus responding directly?
- Can forcing warrant checking through structured prompts improve LLM reasoning?
- Why do LLMs explain evidence accurately while missing its implications?
- Why do LLMs excel at generation but struggle with evaluation?
- How much of LLM reasoning failure stems from missing knowledge versus signal weighting?
- How can a model explain something correctly yet fail to apply it?
- Why do LLMs fail at counterfactual reasoning despite factual knowledge?
- Can surface-level correctness hide failures in structural learning by LLMs?
- How do LLMs handle false presuppositions embedded in user questions?
- What structural framework prevents LLM explanations from becoming just plausible fiction?
- Does prompting for accuracy actually reduce LLM hallucinations and errors?
- Can users experience the LLM Fallacy even when AI outputs are completely accurate?
- How do LLM explanations diverge from actual internal reasoning?
- What structural barriers prevent LLMs from making evaluative judgments about writing?
- Do LLMs understand implicit warrants in reasoning chains?
- How do structured prompts force LLMs to check for contradictions in evidence?
- What happens when we treat LLM outputs as sampled rather than stored?
- Can critique-only calls in LLMs exploit a measurable gap between generation and evaluation?
- Why does analytical depth demand trigger fabrication over transparent uncertainty?
- Why do LLMs fall for and deploy logical fallacies with equal confidence?
- Can prompting a deceptive role change how an LLM tailors its lies?
- Why can LLMs interpret formal logic better than they generate it?
- Can training procedures fix LLM accommodation of false presuppositions?
- Why do LLMs fail when asked to use counter-commonsense rules explicitly?
- Does exposure to more domain-specific examples reduce LLM overconfidence?
- Can LLM-generated descriptions of schemes outperform formal dictionary definitions for prompting?
- Can alignment techniques make LLM explainers match their recommendation behavior?
- Can LLMs explain concepts correctly while failing to use them?
- Can researchers prevent their expectations from shaping LLM outputs?
- Why do users systematically overrely on confident LLM outputs across languages?
- Why do LLMs systematically prefer text from their own family?
- How does token-by-token probability differ from exploring competing rhetorical positions?
- Why do LLM explanations cite similarity and diversity more as options increase?
- Does reasoning happen in hidden space or in generated tokens?
- How do completeness scaffolds force explicit step-by-step derivation?
- Why does probability of text completion not equal knowledge value?
- Where do LLMs succeed at generation but struggle with evaluation?
- Why do entities trigger memorized propositions instead of enabling reasoning?
- Do scheme critical questions work better than direct scheme classification prompts?
- How does era sensitivity in legal cases compound with context length failures?
- How do years of A/B testing compare to one-shot LLM content generation?
- Can auditing LLM performance on complex inputs improve NLP pipeline reliability?
- Can prompted or fine-tuned models generate genuine narrative ambiguity?
- How do embedding contexts like presupposition triggers affect LLM entailment reasoning?
- How much does question framing affect LLM accuracy on knowledge tasks?
- What causes LLMs to ignore unstated constraints they know about?
- What prompting strategies most effectively boost long-context LLM performance on retrieval?
- How can we verify outputs from systems that generate without grounding?
- Can irrelevant information reliably expose the limits of LLM reasoning?
- Why can't LLMs reason from first principles or initial commitments?
- Do longer prediction horizons systematically degrade LLM forecasting accuracy?
- Which use cases can tolerate unverified LLM outputs without external verification?
- Can you detect LLM arguments by measuring convergence with the original post?
- Why does regenerating LLM responses produce different but equally valid answers?
- Why do LLM explanations feel authoritative even when alignment with the model fails?
- What constrains LLM generation beyond default politeness in review contexts?
- Can LLMs propose pivots that change what counts as background context?
- Why do experts experiencing the LLM Fallacy fail to develop custodian skills?
- How does the inability to manage ambiguity undermine literary analysis tasks?
- How does the absence of evaluative stance appear in LLM academic writing?
- How do you partition LLM experts by domain versus by time?
- How does smooth probabilistic flow differ from turbulent rhetorical exploration?
- Could real-time search systems avoid era sensitivity in legal reasoning?
- Can domain pretraining on historical legal corpora reduce era sensitivity?
- Why do LLM personas struggle with specificity in specialized domains like law?
- Why does LLM fluency create false perceptions of professional standing and expertise?
- Why do review corpora contain biases that affect generated comparisons?
- Why do LLM stories over-explain themes and favor single-track plots?
- Can knowledge density explain why LLM writing feels coherent but fatiguing?
- How does the LLM Fallacy differ from automation bias and cognitive offloading?
- What role do model-based critics play in validating LLM plans?
- How does the LLM Fallacy prevent users from noticing cognitive debt accumulating?
- What happens when experts prompt using their own technical register?
- Why does ad-hoc prompt engineering violate scientific method standards?
- How do held-out gates compare as defenses when the proposer is an LLM?