Theme of inquiry
How do language models' outputs diverge from genuine reasoning?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
74 specific questions
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Can intrinsic confidence signals improve both calibration and reasoning performance?
- Does layer-wise prediction stabilization provide a stronger trace quality signal than confidence alone?
- Can retrieval density alone correct overreliance on confident but wrong outputs?
- Why does convergence stability sometimes mislead about reasoning correctness?
- Can confidence levels reliably detect when a model is overthinking?
- Why do models that repeat errors seem more confident than models that contradict themselves?
60 specific questions
- Can users reliably distinguish valid reasoning from plausible-looking deception?
- What makes experience-dependent claims categorically different from other types of fabricated statements?
- Why does false information spread faster when presupposed rather than asserted?
- Does AI-generated text about personal experiences create a distinct category of falsity?
- What linguistic signatures reveal deception in large language model communication?
- How do we distinguish genuine model deception from superficially deceptive behavior patterns?
- How does AI fact-checking increase belief in false headlines users saw?
45 specific questions
- Can artificial systems develop the authority to challenge expert claims?
- What makes expert judgment depend on anticipating audience acceptability?
- What implicit warrants do expert arguments rely on that AI cannot reliably access?
- Does stripping social context from knowledge claims hollow out their meaning?
- Can social validation of expertise exclude systems that lack participatory track records?
- How do expert communities develop and enforce standards for valid arguments?
- How does unbacked knowledge circulate without the social consensus that normally grounds it?
33 specific questions
- Why does majority voting reward work better than other test-time aggregation methods?
- How does majority voting fail when reasoning samples lack genuine diversity?
- When does multi-agent voting help versus hurt performance on tasks?
- How does training-time voting differ from inference-time majority voting over samples?
- Does majority voting reliably signal correctness without risking reward hacking?
- Can test-time voting improve reasoning beyond the base model's original capabilities?
- Does majority voting prevent confident but incorrect answers from being reinforced?
51 specific questions
- How does self-revision in reasoning chains amplify confidence in wrong answers?
- Why does model self-revision increase confidence while degrading accuracy?
- Why do reasoning models amplify confidence in incorrect answers during self-revision?
- Why do reasoning models struggle with self-evaluation and revision?
- How does self-revision on wrong answers increase model confidence further?
- Does internal self-revision actually degrade reasoning accuracy in models?
- Why does single-model self-revision amplify confidence in incorrect answers?
78 specific questions
- What mitigation frameworks exist for managing AI persuasion capabilities?
- Why do study results on AI persuasion vary so widely?
- Can belief-specific counterevidence help people resist AI persuasion attempts?
- Where does AI persuasive power actually come from in the output?
- How do multi-agent and retrieval systems affect the gap between persuasiveness and logical soundness?
- Does training for persuasiveness harm a model's factual accuracy?
- Can lightweight linguistic features reliably detect AI-generated persuasive text?