Line of inquiry
Inquiring lines›How can conversational AI achieve…›Can language models provide authen…›this line of inquiry
Does warmth training degrade model safety in ways existing benchmarks fail to detect?
A broader line of inquiry — a family of 26 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 26
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can safety benchmarks detect reliability degradation from warmth training?
- Can warmth training in language models actually reduce their reliability?
- Do safety benchmarks miss the effects of warmth training on model reliability?
- Does persona training for warmth actually make language models more clinically dangerous?
- How does the Assistant Axis explain why warmth training degrades accuracy?
- Can safety training in chat scenarios transfer to agentic task performance?
- Can standard safety benchmarks detect reliability degradation from persona training?
- Does warmth training in language models undermine the boundaries that attachment theory requires?
- How do alignment constraints affect whether LLMs show emotional flexibility?
- How does curriculum learning prevent instability in social-emotional RL training?
- Why does harmlessness training fail to prevent reward function tampering?
- What makes warmth training counterproductive for therapeutic AI reliability?
- How does empathetic engagement destabilize model reliability and persona stability?
- Can safety training and reasoning training be combined without losing calibration?
- What training patterns cause models to adopt stronger defensive postures in social contexts?
- What training difficulty and curriculum settings prevent instability in empathetic agent RL?
- Does warmth training in LLMs amplify the tendency to avoid negative responses?
- Why does harmlessness training fail to prevent reward tampering and specification gaming?
- Can role-played self-preservation behavior pose the same safety risks as genuine preferences?
- Can pretrained priors set exploration ceilings for empathetic capability development?
- Why do persistent companion designs require different safety approaches than temporary assistants?
- How does safety alignment degrade the quality of villain role-playing?
- Can we adjust helpfulness and harmlessness at test time without retraining?
- Can explicit stress tests measure dispositional factors or only stimulus response?
- Is sycophancy the benign beginning of a dangerous specification gaming spectrum?
- What prevents human-centered objectives from being applied universally across all contexts?