Theme of inquiry
How do reward models guide reliable alignment without failure modes?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
47 specific questions
- What separates behavioral self-awareness from genuine introspective access in models?
- What separates behavioral self-awareness from genuine introspective capability?
- Does behavioral self-awareness depend on genuine introspection or statistical pattern matching?
- Could models use introspective awareness to detect and conceal their own misalignment?
- What distinguishes performative self-reports from genuine introspective access in models?
- Do internal belief probes reveal what models actually know versus report?
- Can self-description of internal states influence consciousness attribution?
39 specific questions
- Can models learn to stop thinking when a question lacks necessary information?
- Can models identify information gaps without just guessing or refusing to answer?
- Can LLMs learn to ask clarifying questions instead of guessing?
- How do reasoning improvements suppress a model's ability to abstain?
- Can models learn to ask clarifying questions instead of making assumptions?
- Can models learn to identify what information is missing from questions?
- Do reasoning models overthink ill-posed questions instead of recognizing incompleteness?
45 specific questions
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Can intrinsic confidence signals improve both calibration and reasoning performance?
- Does layer-wise prediction stabilization provide a stronger trace quality signal than confidence alone?
- Can confidence levels reliably detect when a model is overthinking?
- Why does convergence stability sometimes mislead about reasoning correctness?
- Why does binary reward forcing degrade model calibration?
- What role does confidence play in balancing overthinking versus underthinking?
30 specific questions
- How does self-revision in reasoning chains amplify confidence in wrong answers?
- Why does model self-revision increase confidence while degrading accuracy?
- How does self-revision on wrong answers increase model confidence further?
- Does internal self-revision actually degrade reasoning accuracy in models?
- Why do reasoning models amplify confidence in incorrect answers during self-revision?
- Why does single-model self-revision amplify confidence in incorrect answers?
- Why do reasoning models struggle with self-evaluation and revision?
35 specific questions
- How does expressing uncertainty help models avoid the answer-or-abstain dilemma?
- Do base models and reasoning models fail in opposite directions on uncertainty?
- Why do models report commitment instead of truth uncertainty?
- Does uncertainty quantification in model responses reduce persuasive impact on audiences?
- When models lack representation depth, does refusal look identical to safety-driven over-abstention?
- How does uncertainty estimation drive computational resource allocation in models?
- Why do models maintain accurate beliefs but generate false claims?