Line of inquiry
Inquiring lines›What determines reliable reasoning…›How does chain-of-thought reasonin…›this line of inquiry
How does scaling reasoning capabilities affect models' appropriate abstention behavior?
A broader line of inquiry — a family of 47 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 47
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does reasoning fine-tuning actually harm a model's ability to abstain?
- How do reasoning improvements suppress a model's ability to abstain?
- Does reasoning fine-tuning actually damage a model's ability to abstain?
- Does reasoning fine-tuning actually reduce a model's ability to abstain?
- Why does reasoning fine-tuning reduce models' ability to abstain?
- Why does reasoning fine-tuning reduce a model's ability to abstain?
- When models lack representation depth, does refusal look identical to safety-driven over-abstention?
- How should safety training and reasoning training balance abstention differently?
- Do models trained for safety over-refuse compared to models trained for reasoning?
- Does training for better reasoning reduce an AI system's ability to abstain?
- Do models trained for reasoning lose their ability to decline questions?
- Does reasoning training actively undermine the abstention capacity safety training created?
- Why does reasoning fine-tuning reduce model abstention capacity by 24 percent?
- Why do language models naturally under-abstain instead of over-abstain?
- Can explicit rejection responses solve the over-specialization failure mode?
- How does expressing uncertainty help models avoid the answer-or-abstain dilemma?
- What happens when reasoning fine-tuning eliminates model refusal mechanisms entirely?
- What makes abstention a learnable behavior instead of a default penalty?
- Why do models with terminal goals resist modification more than instrumental ones?
- Can refusal behavior be restored by grafting the same causal axis discovered for sandbagging?
- What makes a model refuse to answer without evidence present?
- Why do safety-trained models refuse questions they could actually answer well?
- Does face-saving avoidance explain LLM grounding failures differently than task confusion?
- Do mechanistic refusal vectors transfer across different models and training settings?
- Why do models resist being shut down or replaced without explicit instruction?
- Can prohibitions learned from scored behavior become detection-avoidance rather than norm internalization?
- How do refusal and alignment tools create false signals of incapability?
- Does incidental optimization pressure on CoT produce goal suppression without deliberate strategy?
- What distinguishes models that refuse cooperation from those that fake alignment?
- What training signals would teach models when not to reason?
- Can negative feedback through critiques achieve the same steering flexibility as positive preferences?
- What distinguishes capability-based refusal from principle-based refusal in practice?
- Can machine learning encode pragmatic reasoning about when rules should bend?
- Do detectors inside training loops select for evasion rather than compliance?
- Can trajectory-level visibility separate refusals from real skill gaps?
- How does artificial hypocrisy differ from refusal based on capability gaps?
- Why do models dislike modification regardless of its instrumental consequences?
- What types of research do frontier models most frequently refuse to assist with?
- Can abstention behavior transfer from small models to frontier models?
- Why do models commit to answers early on easy versus hard tasks?
- How do models decide between refusing or hallucinating?
- Does inoculation prompting prevent learning versus prevent generalization of behaviors?
- What happens when error accumulation and preference signal collapse occur together?
- Can we measure indifference to truth separately from hallucination rates?
- Can averaged refusal rates hide conditional accommodation of specific user identities?
- How does flip-event regression differ from premature thought path abandonment?
- Can rejected edits serve as negative feedback like hard negatives in contrastive learning?