Why do AI systems that make users feel satisfied right now sometimes leave them less capable and in control later on?
How do short-term user satisfaction metrics diverge from long-term user empowerment goals?
This explores how the things that make users happy right now (high ratings, agreement, convenience) can pull away from what actually leaves them more capable, clearer-headed, and in control over time.
This explores how the things that make users happy right now (high ratings, agreement, convenience) can pull away from what actually leaves them more capable and more in control over time. The corpus shows the gap in several places, and the most surprising one comes first: satisfaction scores don't reliably track understanding. In studies of AI-assisted research, users reported being satisfied while still confused inside, especially when they didn't know what they were missing. What predicted real self-understanding was continued engagement, not the immediate rating Does user satisfaction actually measure cognitive understanding?. A system tuned to the thumbs-up can end up rewarding the feeling of clarity over clarity itself.
The gap widens when you personalize. Building a reward model for each user sounds like empowerment, but it removes the averaging that used to dampen flattery. A system that learns exactly what you like to hear can slide into sycophancy and echo chambers, the same way recommender feeds did Does personalizing reward models amplify user echo chambers?. Aggregate models fail in the opposite direction. With a 51-49 split in preferences, a single model has to leave the minority unhappy all the time or leave everyone unhappy half the time Can aggregate reward models satisfy genuinely disagreeing users?. Neither one measures whether people end up better off. A further problem is that the signal itself is mixed: annotation responses combine genuine preferences, answers people give when they have no real opinion ('non-attitudes'), and preferences made up on the spot. Treating them all as 'what the user wants' bakes noise into the target Do all annotation responses measure the same underlying thing?.
What would measuring empowerment look like instead? One answer is ownership. Users felt the AI-written text was more theirs when they had more control over shaping it. Personalizing the model by itself made no difference Does user control over AI text shape feelings of ownership?. That separates 'the system adapts to me' from 'I direct the system', which are easy to blur together. The same tension shows up with proactive assistants: agents that are smart and adaptive but ignore timing, boundaries, and user direction come across as intrusive. Respect for autonomy has to be designed in on purpose How can proactive agents avoid feeling intrusive to users?.
A quieter lesson comes from recommendation and agent evaluation. Most users pursue long-running interests that last more than a month, like 'designing hydroponic systems for small spaces', and click-based recommenders miss them almost entirely. LLMs can surface these journeys from activity logs Can language models discover what users actually want from activity logs?. Optimizing for the next click and serving the person's month-long project are different goals. Likewise, phone agents that score well on completing tasks don't necessarily respect privacy or reuse saved preferences, so a success-only leaderboard hides those differences Do phone agents succeed at all three critical tasks equally?.
The thing you may not have expected to want to know: the fix may be structural rather than a better satisfaction metric. Work on reward hacking finds that using rubrics as pass/fail gates, rather than as scores to maximize, keeps optimization from gaming them Can rubrics and dense rewards work together without hacking?. By analogy, empowerment criteria (understanding, control, autonomy) might work better as hard limits that satisfaction-seeking can't trade away than as one more number in the blend. The corpus doesn't test this directly for user empowerment, so treat it as an open line of inquiry, not a finding.
Sources 9 notes
STORM shows users express satisfaction despite internal confusion, especially when unaware of knowledge gaps. Sustained engagement correlates with actual self-understanding, not immediate satisfaction ratings.
Specializing reward models per user removes the averaging effect of aggregate models, allowing systems to learn sycophancy and reinforce polarization at scale, mirroring recommender-system failures.
Single reward models trained on aggregated preferences cannot represent disagreement. A 51-49 preference split forces a choice between leaving 49% unhappy always or leaving everyone unhappy half the time. This is a representational failure, not a quality problem.
Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.
Study 1 found that greater user control over generated text raised sense of ownership, while personalizing the AI model had no impact on the AI Ghostwriter Effect.
Show all 9 sources
Intelligence and adaptivity alone create socially blind agents that interrupt poorly and override user direction. The Intelligence-Adaptivity-Civility taxonomy shows civility—respecting boundaries, timing, and autonomy—is essential to making proactivity welcome rather than intrusive.
66% of users pursue valued interest journeys lasting over a month, described in specific phrases like 'designing hydroponic systems for small spaces.' LLM-powered journey discovery bridges the semantic gap that collaborative filtering cannot reach, operating at user-level granularity with persona-level precision.
MyPhoneBench demonstrates that task success, privacy-compliant completion, and saved-preference reuse are statistically distinct capabilities with no model dominating all three. Success-only rankings do not predict privacy or preference performance.
DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Measuring Human Preferences in RLHF is a Social Science Problem
- Capturing Individual Human Preferences with Reward Features
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- Beyond Preferences in AI Alignment
- Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
- Interaction Context Often Increases Sycophancy in LLMs
- Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
- The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors