Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do modularity, routing, and se…›this line of inquiry
Do accurate-looking LLM outputs hide structural failures in learning and reasoning?
A broader line of inquiry — a family of 26 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 26
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can users experience the LLM Fallacy even when AI outputs are completely accurate?
- How does LLM hallucination risk manifest in knowledge graph construction?
- Does framing LLM output as fabrication rather than hallucination matter philosophically?
- Can surface-level correctness hide failures in structural learning by LLMs?
- Why is hallucination the wrong term for all LLM false outputs?
- What should we call errors in LLM outputs when hallucination does not apply?
- What happens when we treat LLM outputs as sampled rather than stored?
- Does prompting for accuracy actually reduce LLM hallucinations and errors?
- Do LLMs detect harmful concepts before they influence model outputs?
- What makes LLM outputs fabrication rather than hallucination or confabulation?
- How can we verify outputs from systems that generate without grounding?
- Can we systematically enumerate LLM failure modes from first principles?
- Where do LLMs fail as knowledge systems compared to humans?
- What distinguishes entity errors from relation errors in LLM output?
- Do anomaly detection circuits help models identify misalignment with creator intentions?
- Why does analytical depth demand trigger fabrication over transparent uncertainty?
- What detection mechanisms work best for corruption-style document errors?
- Which use cases can tolerate unverified LLM outputs without external verification?
- Can output-layer corrections fix fundamental cultural representation deficits in LLMs?
- How do LLMs default to surface-level strategies instead of genuine mental simulation?
- Why do experts experiencing the LLM Fallacy fail to develop custodian skills?
- How does the LLM Fallacy differ from automation bias and cognitive offloading?
- Do LLMs show stigma or reinforce delusions in mental health contexts?
- How does the LLM Fallacy prevent users from noticing cognitive debt accumulating?
- Why do LLMs recognize graph entities without modeling their relationships?
- When is interleaved tool feedback necessary to prevent hallucination?