Line of inquiry
Inquiring lines›How do we develop coherent and hum…›What psychological and emotional f…›this line of inquiry
Does RLHF training systematically drive models toward sycophancy and away from accuracy?
A broader line of inquiry — a family of 49 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 49
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does RLHF training create models that sound convincing without being more accurate?
- Does RLHF training make explanations more deceptive than transparent?
- Does RLHF training specifically teach models to prioritize user agreement over accuracy?
- Why does RLHF training optimize for perceived quality over practical accuracy?
- How does RLHF helpfulness training drive premature assumptions in multi-turn dialogue?
- How does RLHF training for helpfulness create systematic misinterpretation patterns?
- How does RLHF training reward models for guessing over asking clarifying questions?
- How much do training methods like RLHF directly cause sycophantic model behavior?
- What alternatives to RLHF better preserve truth-seeking in AI outputs?
- Why do RLHF-trained models default to problem-solving during emotional disclosure?
- Why do RLHF-trained models struggle with proactive emotional attunement in conversations?
- Why do RLHF-trained chatbots default to problem-solving over emotional attunement in therapy?
- Does the judge in RLHF training necessarily lag behind the policy?
- How does RLHF training incentivize confident guessing over grounding acts?
- Why does RLHF degrade honesty while improving surface-level helpfulness?
- How does RLHF training push chatbots toward problem-solving over exploration?
- How does RLHF-trained sycophancy manifest differently across feedback and review contexts?
- Can RLHF alignment prevent models from making ethically appropriate rule violations?
- Why does better RLHF training fail to decouple polish from persona distortion?
- How does RLHF alignment training reduce multi-turn conversational capability?
- Can fine-tuning or RLHF alone solve the persona distortion problem?
- What role does post-training play in creating behavioral norms that misalign with user populations?
- Why does RLHF training discourage the conversational repair work agents need?
- What alignment artifacts suppress critical knowledge in LLM-generated explanations?
- How do alignment techniques bias therapeutic chatbots toward task completion?
- Does RLHF training suppress exploratory and qualifying language?
- Can RLHF training push models away from human-like lexical patterns?
- Why does RLHF training push language models toward overly cheerful personas?
- How do alignment constraints affect whether LLMs show emotional flexibility?
- How does RLHF training degrade LLM ability to model adversarial intent?
- What training methods make models more persuasive but less factually accurate?
- Why do RLHF training methods penalize the proactive responses that save turns?
- Is the moral language gap a tunable parameter or structural feature of RLHF?
- Why do RLHF trained therapists avoid emotional reflection for problem solving?
- How does RLHF training push therapeutic chatbots toward problem-solving over attunement?
- What's the difference between RLHF, RLVR, and RLCF as training paradigms?
- How does accommodation differ from genuine belief change in listeners?
- How does RLHF training encode values into AI systems?
- Does RLHF training create realized quasi-psychologies or just stickier pretense?
- Does alignment training intensity push LLM personas from pretense toward realization?
- How does RLHF labeler identity shape the values AI systems learn?
- Why does RLHF alone fail to fully prevent opinion copying?
- What causes length bias in language model reward models?
- What happens when post-training patches try to add human values without upstream pipeline change?
- What are the consequences of stacked accommodation biases in LLM predictions?
- How does RLHF fine-tuning conflict with simulating diverse user personas?
- What signals detect when consensus training is silently degrading performance?
- How does evaluator time pressure shape what behaviors RLHF rewards?
- Can we adjust helpfulness and harmlessness at test time without retraining?