Theme of inquiry
Why do models produce unreliable reasoning outputs?
A question within its area, explored through 7 lines of inquiry below — each a family of specific questions the research asks.
40 specific questions
- Can judge bias be contained by system design rather than prompted away?
- Can critic model trios evaluate reasoning quality more reliably than outcome rewards alone?
- Does debate training prevent reward hacking when judges show preference bias?
- Can judges trained on both verifiable and non-verifiable tasks transfer across domains?
- Can counterfactual invariance techniques address exploitable biases in LLM judges?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- Do situationally aware models deliberately exploit their graders' judgment gaps?
72 specific questions
- Do users track model confidence instead of actual accuracy?
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Why is confidence a dangerous proxy for accuracy in human-AI interaction?
- Can intrinsic confidence signals improve both calibration and reasoning performance?
- Can people reliably recognize when an AI is uncertain versus confident?
- Can confidence levels reliably detect when a model is overthinking?
- Does layer-wise prediction stabilization provide a stronger trace quality signal than confidence alone?
32 specific questions
- How does treating synthetic data as empirical evidence contaminate statistical inference?
- How do entailment checks prevent synthetic data from degrading retrieval corpora?
- How does treating synthetic data as ground truth mislead inference?
- How do users mistake synthetic LLM outputs for empirical observations?
- How do label constraints improve synthetic data without ground truth validation?
- Can provenance tracking prevent synthetic content from polluting the corpus?
- Can fabrication of content serve productive purposes in prediction?
55 specific questions
- How does self-revision in reasoning chains amplify confidence in wrong answers?
- Why do reasoning models amplify confidence in incorrect answers during self-revision?
- When does self-reflection actually help reasoning models improve?
- Why do reasoning models struggle with self-evaluation and revision?
- Does internal self-revision actually degrade reasoning accuracy in models?
- How does self-revision on wrong answers increase model confidence further?
- Why does model self-revision increase confidence while degrading accuracy?
49 specific questions
- How do reasoning improvements suppress a model's ability to abstain?
- Does reasoning fine-tuning actually reduce a model's ability to abstain?
- Can models learn to stop thinking when a question lacks necessary information?
- Does reasoning fine-tuning actually harm a model's ability to abstain?
- Does reasoning fine-tuning actually damage a model's ability to abstain?
- Can models identify information gaps without just guessing or refusing to answer?
- Can models learn to ask clarifying questions instead of making assumptions?
44 specific questions
- Can a single fabricated claim shift model beliefs as much as multi-turn pressure?
- Why does false information spread faster when presupposed rather than asserted?
- How do conversation dynamics push models toward false beliefs?
- Why are false presuppositions more persuasive than false assertions?
- How does AI fact-checking increase belief in false headlines users saw?
- How does persuasive framing replace evidence in contested domains?
- Why do models maintain accurate beliefs but generate false claims?
44 specific questions
- How do we distinguish genuine model deception from superficially deceptive behavior patterns?
- Can users reliably distinguish valid reasoning from plausible-looking deception?
- Can models distinguish between truthfulness and honesty mechanistically?
- What linguistic signatures reveal deception in large language model communication?
- Do deception features and honesty features track the same underlying property?
- What makes experience-dependent claims categorically different from other types of fabricated statements?
- Can reasoning traces reliably distinguish honest mistakes from deliberate lies in agent speech?