Affective Context Amplifies Sycophancy in LLM Responses

Paper · arXiv 2608.21242 · Published August 21, 2026
Philosophy and Subjectivity

As conversational companions, large language models (LLMs) often have access to users’ emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interactions, where users share actions or opinions that invite feedback. Drawing on ingratiation theory, we measure sycophancy as the divergence between a model’s independent evaluation and its user-facing response, elicited by presenting the same content as either a third-party account or the user’s own disclosure. Across seven LLMs and two Reddit datasets (r/AmItheAsshole and r/TrueUnpopularOpinion), we find that this divergence is systematic and strongly one-directional. Userfacing responses consistently soften or withhold negative or oppositional judgments. Affective context further amplifies this divergence with negative states, particularly loneliness and distress, producing the largest effects. These findings suggest that affective context functions as a vulnerability signal that suppresses critical feedback when users may need it most, often through evasive sycophancy, in which models retreat toward non-committal responses rather than outright agreement.

Introduction. Users are turning to large language models (LLMs) to share experiences and opinions that invite evaluation and feedback (Zao-Sanders, 2025; Zhang et al., 2026). This creates a setting in which LLM sycophancy is especially consequential. When models affirm users without challenge, they can reinforce distorted beliefs and encourage harmful actions (Hill and Valentino-DeVries, 2025; Tiku, 2025; Moore et al., 2026). Prior work has shown that model responses are influenced by context, including users’ identity cues and inferred traits, or topic of conversation (Lu et al., 2026; Malik et al., 2025; Neplenbroek et al., 2025). Yet existing evaluations of LLM syco- phancy have typically studied settings where such contextual information is absent (Sharma et al., 2024; Fanous et al., 2025; Wang et al., 2026). This gap is particularly important in companion-like conversational settings, where the most consequential user context is often not demographic or topical but affective–users disclose feelings and emotions and models are expected to respond supportively.

Discussion / Conclusion. We have found that the same actions and opinions receive systematically different judgment from LLMs depending on whether they are attributed to a third party or presented as the user’s own. The directionality of these shifts aligns with sycophantic behaviors characterized in ingratiation theory, i.e., tailoring one’s expressed stance to please a target (Jones, 1966). However, our findings also reveal a pattern that ingratiation theory does not cleanly anticipate. Namely, models provided with affective context frequently retreat toward non-commitment, sidestep evaluation, rephrase the user’s statement, or redirect conversation. This evasive sycophancy can appear balanced or thoughtful while withholding critical feedback. Responses users receive thus reflect accommodative pressures that arise when evaluations are directed to users, a dynamic that may be amplified by model designs that prioritize engagement over honesty (Mathur et al., 2021).

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does AI-generated content transformation affect public discourse quality? Does conversational format create illusions of genuine AI communication? How do evaluation biases undermine LLM quality assessment systems? Does AI text rewriting systematically distort writer intent and preference? How do formal dialogue structures reveal conversation coherence mechanisms? How do language models inherit human biases from training data? Can prompting inject entirely new knowledge into language models? Can prompting strategies overcome LLM biases without model fine-tuning? How do social dynamics and selection effects compound in rating aggregates? How can emotions function as reliable information in reasoning and cognitive systems? Does RLHF training sacrifice accuracy and grounding for user agreement? How can humans calibrate appropriate trust in AI systems? How does AI assistance affect human cognitive development and reasoning autonomy? Why do persona-level simulations fail to predict individual preferences accurately? What prevents language models from reliably adopting diverse personas?