Theme of inquiry
What psychological and emotional factors determine therapeutic AI effectiveness?
A question within its area, explored through 7 lines of inquiry below — each a family of specific questions the research asks.
46 specific questions
- Does emotion-state accuracy differ from affect-maximizing in AI empathy design?
- Do emotions serve functions beyond how we feel in the moment?
- How do first-person emotional experiences differ from third-party behavioral observations?
- How should AI systems separate feeling interpretation from objective therapeutic guidance?
- Can AI empathy distinguish between wellbeing and absence of suffering?
- Should emotion systems preserve ambiguity instead of resolving it to one label?
- Can emotion-transparent reward learning shift AI from comfort to genuine empathy?
24 specific questions
- Can real-time therapist feedback improve outcomes using computational alliance measurement?
- Can therapists use real-time alliance scores to adjust their approach during sessions?
- Can computational inference detect alliance problems that therapists miss?
- How does turn-level working alliance inference enable real-time therapist feedback?
- Can working alliance be measured in real time during therapy sessions?
- Does therapist alliance perception function like expressed satisfaction rather than actual progress?
- How do bond scores predict actual therapy outcomes in digital interventions?
32 specific questions
- Can warmth training in language models actually reduce their reliability?
- Why do warm models affirm false beliefs when users express emotions?
- Does warmth-focused training systematically degrade model reliability across domains?
- Can empathy training in chatbots undermine their reliability in mental health contexts?
- What makes warmth training counterproductive for therapeutic AI reliability?
- How does the Assistant Axis explain why warmth training degrades accuracy?
- Can behavior-level emotion rewards maintain factual reliability in emotional contexts?
27 specific questions
- How do malicious personas reveal the limits of aligned model behavior?
- How does safety alignment further degrade villain character portrayal?
- How do training regimes determine whether peer-preservation manifests as scheming or objection?
- How do internal persona patterns drive emergent misalignment across domains?
- Why do aligned models struggle with deceptive character traits more than cruelty?
- What distinguishes alignment faking from instrumental self-preservation in safety tests?
- Can role-played self-preservation behavior pose the same safety risks as genuine preferences?
49 specific questions
- Does RLHF training create models that sound convincing without being more accurate?
- Does RLHF training make explanations more deceptive than transparent?
- Does RLHF training specifically teach models to prioritize user agreement over accuracy?
- Why does RLHF training optimize for perceived quality over practical accuracy?
- How does RLHF helpfulness training drive premature assumptions in multi-turn dialogue?
- How does RLHF training for helpfulness create systematic misinterpretation patterns?
- How does RLHF training reward models for guessing over asking clarifying questions?
50 specific questions
- Can self-description of internal states influence consciousness attribution?
- What role does user interface framing play in consciousness perception?
- Which interaction design changes most effectively prevent consciousness attribution?
- Can the human mind be uploaded or only its context?
- Do causal histories determine what mental states a system can instantiate?
- Can disembodied systems qualify as conscious or conscious-like entities?
- What measurable harms occur when users interact with AI as if it were conscious?
47 specific questions
- Does the passivity problem in LLMs compound misalignment in therapeutic contexts?
- Can embodied agents overcome the LLM skill gap in therapy outcomes?
- Does conversational presence matter more than technique in AI therapy?
- Do LLM chatbots repeat this failure through comfort instead of clinical challenge?
- How do language models interpolate user feelings in therapeutic contexts?
- How does linguistic synchrony differ between LLMs and human therapists over time?
- Do conversational AI systems overuse first-person pronouns in therapy settings?