Line of inquiry
Inquiring lines›Why are language models fragile de…›How do learned model representatio…›this line of inquiry
Are language model reasoning explanations faithful to their actual thinking?
A broader line of inquiry — a family of 66 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 66
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can language models accurately evaluate the quality of their own reasoning?
- Can we distinguish between semantic and symbolic reasoning in language models?
- Why do language models produce unfaithful chain of thought explanations?
- Why do language models generate reasoning tokens after internally deciding the answer?
- Can language models correct false assumptions or only reinforce them?
- Does more thinking always improve language model accuracy?
- How does hidden processing in language models prevent accurate self-assessment?
- Can models maintain reasoning-output coupling while improving domain accuracy?
- Why does answer-confirmation bias emerge in language model reasoning?
- What structural properties of language models make fabrication inevitable?
- What makes truthfulness and honesty mechanistically different in language models?
- Does premature confidence signal flawed reasoning in language models?
- What distinguishes genuine understanding from correct output without coherent principles?
- Why do generative and discriminative language model procedures disagree?
- How do mechanistic interpretability methods surface what models represent internally?
- What distinguishes real understanding from superficial pattern matching?
- What explains the gap between perplexity performance and actual reasoning capability?
- Why do language models fail at grounding and inference?
- Why does search-augmented generation still not solve the verification problem?
- How can we measure whether an agent reasons correctly rather than just sounds plausible?
- Why do explicit linguistic markers override semantic computation in models?
- Can correct model outputs prove that semantic meaning rather than surface patterns drove the response?
- Can language models distinguish between novel insight and unjustified conceptual blending?
- How does the generation-verification gap prevent language models from improving themselves?
- How does evaluation setting affect measured reasoning capabilities in language models?
- What reveals the epistemic limits of language models?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Why do language models naturally under-abstain instead of over-abstain?
- Can language models distinguish explicit from implicit discourse relations?
- Why do NLP models fail at recognizing multiple valid interpretations?
- Why does evaluating errors teach more than imitating correct responses?
- Can contamination-free evaluation distinguish between memorization and genuine prediction ability?
- Do language models favor outputs from their own model family?
- Do dialogue agents have authentic voice agency or beliefs of their own?
- Can models detect false presuppositions when they actually possess the knowledge?
- How does linguistic calibration differ from token probability calibration?
- Can external classifiers reliably decide when a model should reason?
- Why does consistency training make models resistant to prompt perturbations?
- How does tool integration leverage comprehension without demanding perfect generation?
- What is the difference between a truthful answer and an honest one?
- Why do benchmarks measuring string quality fail to capture communicative success?
- How do semantic and symbolic reasoning capabilities differ in language models?
- Why are false presuppositions harder to spot when they sound plausible?
- What role do humans play in converting language model outputs into meaningful events?
- Can you separate grammatical competence from rhetorical commitment in language systems?
- Why do language models produce verbose reasoning when asked to think step by step?
- When do language models first develop self-preservation preferences?
- What design changes could make constraint inference more reliable without explicit cuing?
- What implicit premises do language models skip even with correct surface reasoning?
- What separates pattern matching from genuine language understanding?
- Why do users attribute consciousness to language models in practice?
- How does predictive accuracy on future tokens differ from correctness on labeled answers?
- What makes a problem instance unfamiliar to a language model?
- What distinguishes character simulation from authentic voice in language model outputs?
- Why can language models detect author style without understanding why it matters?
- What's the difference between language generation and human-to-human communication?
- Why do language models fail at pronouns across distant segments?
- Why do different language models independently converge toward similar outputs in open-ended generation?
- What emerges in large language models that makes explicit value modeling necessary?
- What makes a synthetic belief robust versus generative for downstream learning?
- What makes sincerity impossible without a coherent first-person perspective?
- What architectural changes would let language models develop genuine functional competence?
- What would it mean for a language model to canvas counterpositions?
- Why do fluent model outputs resist challenge despite containing injected content?
- What is the comprehension-generation asymmetry in language models?
- Do models learn different sophistry strategies for QA versus code generation?