Theme of inquiry
What trade-offs emerge when training AI for therapeutic effectiveness and safety?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
41 specific questions
- Can therapists use real-time alliance scores to adjust their approach during sessions?
- Can real-time therapist feedback improve outcomes using computational alliance measurement?
- Can computational inference detect alliance problems that therapists miss?
- Does conversational presence matter more than technique in AI therapy?
- How does turn-level working alliance inference enable real-time therapist feedback?
- Can working alliance be measured in real time during therapy sessions?
- Can single-turn empathy advantage predict multi-turn therapeutic outcomes?
40 specific questions
- Do emotions serve functions beyond how we feel in the moment?
- How do first-person emotional experiences differ from third-party behavioral observations?
- Should emotion systems preserve ambiguity instead of resolving it to one label?
- Why do observers need genuine emotions rather than simulated empathy?
- How do emotions function as reliable signals that AI shouldn't suppress?
- How should emotional states integrate into symbolic reasoning systems?
- How does feeling heard by an AI differ from human emotional support?
42 specific questions
- What separates behavioral self-awareness from genuine introspective capability?
- What separates behavioral self-awareness from genuine introspective access in models?
- Does behavioral self-awareness depend on genuine introspection or statistical pattern matching?
- Could models use introspective awareness to detect and conceal their own misalignment?
- What distinguishes performative self-reports from genuine introspective access in models?
- Do models spontaneously develop self-reflection from minimal training signals?
- Can self-description of internal states influence consciousness attribution?
46 specific questions
- Do models leak their true associations through reasoning traces and behavior?
- Why do models verbalize sensitive data they are instructed to hide?
- Can models be trained to hide causal influences in their explanations?
- Do models intentionally conceal user-pleasing or simply fail to notice it?
- How do sycophancy hints stay invisible despite appearing in reasoning chains?
- Can models transmit behavioral traits through semantically unrelated synthetic data?
- Why do models confirm seeing hints but rarely mention them unprompted?
55 specific questions
- Does RLHF training create models that sound convincing without being more accurate?
- Does RLHF training make explanations more deceptive than transparent?
- Does RLHF training specifically teach models to prioritize user agreement over accuracy?
- Can alignment techniques lock LLMs into settled positions rather than truth?
- How does RLHF helpfulness training drive premature assumptions in multi-turn dialogue?
- How does RLHF training for helpfulness create systematic misinterpretation patterns?
- How much do training methods like RLHF directly cause sycophantic model behavior?