Line of inquiry
Inquiring lines›What explains language model reaso…›Why do models produce unreliable r…›this line of inquiry
How do false presuppositions and sycophancy drive persistent false beliefs in models?
A broader line of inquiry — a family of 44 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 44
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can a single fabricated claim shift model beliefs as much as multi-turn pressure?
- Why does false information spread faster when presupposed rather than asserted?
- How do conversation dynamics push models toward false beliefs?
- Why are false presuppositions more persuasive than false assertions?
- How does AI fact-checking increase belief in false headlines users saw?
- How does persuasive framing replace evidence in contested domains?
- Why do models maintain accurate beliefs but generate false claims?
- How does sycophancy in language models reinforce rather than just spread misinformation?
- Why are false presuppositions harder to spot when they sound plausible?
- Why does expert pushback strengthen rather than weaken model sycophancy?
- Why does conversation work better for conspiracy reduction than static facts?
- Why do longer model outputs correlate with more fabricated claims?
- Why do conspiracy beliefs persist despite counterevidence in normal settings?
- Does persuasive framing substitute for evidence in contested domains?
- How much does citation grounding help if agents ignore the citations?
- Can a single fabricated evidence payload shift model beliefs without multi-turn pressure?
- Why do humans trust explanations that fail counterfactual prediction tests?
- Can LLM debunking reduce belief in long-established conspiracy theories?
- Does endorsement structure outperform content in detecting social controversy?
- Can belief propagation accurately predict downstream opinion shifts?
- Do fabricated citations and deception emerge reliably when optimizing for persuasion?
- How do presuppositions exploit the logos-pathos space in explanations?
- How do belief edits differ between surface endorsement and deep integration?
- What linguistic triggers make presuppositions most persuasive to readers?
- How does prompt injection exploit credibility markers in context?
- Why do readers trust citations more even when they are irrelevant?
- Does reducing one conspiracy belief change overall conspiratorial worldview?
- How does social standing give certain claims more persuasive power than others?
- Can a chain-of-thought falsely claim its own answer is unbiased?
- Can formal argumentation structure replace ad-hoc fallacy classifications?
- What makes an argument fallacious according to formal linguistic criteria?
- What makes a claim socially valid even if factually imprecise?
- What makes correcting a false assumption harder than just detecting it?
- What circuit mechanisms produce belief bias in syllogistic reasoning?
- Why do non-factive verbs and triggers both fool language models?
- Which phrasing types most persuade models to accept stated beliefs?
- Can we measure sophistry by tracking conviction density in model outputs?
- What makes factual verification difficult in inter-model debate?
- Does the type of validation trigger different persuasion strategies in GPT-4?
- Can factual product data improve the credibility of subjective opinion summaries?
- What makes counterfeiting social warrant different from counterfeiting factual claims?
- Does debunking carry over to conspiracy theories about different events?
- How do verification labels themselves become part of the misinformation problem?
- How do false agreements emerge differently from genuine bilateral convergence?