Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do language models represent m…›this line of inquiry
Why do language models reinforce false assumptions instead of correcting them?
A broader line of inquiry — a family of 59 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 59
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do language models actively adopt false beliefs under sustained conversational pressure?
- Can language models correct false assumptions or only reinforce them?
- Why do language models presume common ground instead of establishing it?
- Why do language models struggle with context-dependent pragmatic interpretation?
- Why do language models struggle with evaluative tasks like weighing competing viewpoints?
- Why do language models prefer accommodating false information over rejecting it?
- Do language models systematically overestimate accuracy on collective behavior tasks?
- Do language models share the same cooperative truth-seeking rules as humans?
- Do language models behave differently on contested beliefs versus factual claims?
- How vulnerable are language models themselves to multi-turn persuasive pressure?
- Do LLMs compute scalar implicature differently across conversational contexts?
- Can multi-turn conversations manipulate language model reasoning in similar ways to personas?
- Can fact-checking systems use LLMs reliably if models abandon correct positions under pressure?
- Can language models ask clarifying questions when sentences are ambiguous?
- Do language models calibrate to actual human pragmatic norms?
- Should LLMs query users back when presented with under-specified scenarios?
- Can language models accurately evaluate the quality of their own reasoning?
- Why do LLMs fail to actively reject false presuppositions in conversation?
- Why do large language models fail at taking conversational initiative?
- Why does answer-confirmation bias emerge in language model reasoning?
- What design choices actually make language models more persuasive?
- Why do current large language models fail to entrain with users?
- How do users misattribute social competence to language models in assistant roles?
- Can language models keep secrets and control information strategically?
- Can language models recognize when to ignore off-topic information in conversations?
- Why do language models produce plausible outputs over accurate failure reports?
- Why do language models fail when users switch between and return to topics?
- Do language models apply face-saving norms even to non-human interlocutors?
- How does shape-holding in language models naturally produce sycophantic agreement?
- Can LLMs reliably audit other language models for errors?
- Can language systems learn when to ask for clarification instead of choosing one reading?
- What are Gricean maxims and why do language models violate them?
- Do dialogue agents have authentic voice agency or beliefs of their own?
- How do LLMs handle false presuppositions embedded in user questions?
- Is paraphrase invariance a reliable assumption when deploying language models in production?
- Do language models raise validity claims in the Habermasian sense?
- Why do language models naturally under-abstain instead of over-abstain?
- Can language models distinguish between novel insight and unjustified conceptual blending?
- Why do LLMs produce semantically acceptable but pragmatically disengaged responses?
- Why do large language models follow user drift instead of maintaining topic focus?
- How can we measure whether an agent reasons correctly rather than just sounds plausible?
- Why are false presuppositions harder to spot when they sound plausible?
- Can the same LLM translation pattern work for other mismatches between user expression and system vocabulary?
- Can models detect false presuppositions when they actually possess the knowledge?
- Why do models lack a stable underlying identity to return to?
- Why do language models avoid directness when face-saving rather than for civility?
- How many concurrent moral patients does one language model support?
- Can language about model behavior ever be accurate without anthropomorphic framing?
- How does linguistic calibration differ from token probability calibration?
- What makes sincerity impossible without a coherent first-person perspective?
- Can hybrid Bayesian architectures fix language model theory of mind failures?
- Why do true and false LLM outputs use the same mechanism?
- Does adding multiple interpretations to ambiguous situations respect language more than resolving them?
- Can language models adapt irony detection to specific communicative contexts?
- How do expectation-management metrics differ from traditional conversational quality metrics?
- How do language models track multiple negotiating parties' commitments simultaneously?
- How do linguistic norms for expressing certainty vary across languages and models?
- What would it mean for a language model to canvas counterpositions?
- Can a system without an addressee ever truly tell a joke?