INQUIRING LINE

Does your recommender actually get what you're working toward, or is it just guessing your next click?

Do recommender systems infer journey-level goals or just predict next items?

This explores whether recommender systems understand the longer-running goals people pursue over weeks or months, like learning a hobby or planning a project, or whether they mostly guess the next thing you'll click.


This explores whether recommenders grasp the bigger goals people pursue over time, or mostly predict your next click. The short answer from the corpus is that most systems predict the next item, and the clearest evidence of what they miss comes from work that tries to fill that gap. One study found that 66% of users pursue interest journeys lasting more than a month. These journeys are specific enough to describe in a phrase like 'designing hydroponic systems for small spaces.' Collaborative filtering, which matches you with people who clicked similar things, can't see them. Language models reading activity logs can find and name them Can language models discover what users actually want from activity logs?. The gap isn't missing data. It's that clicks never get turned into a description of what the person is actually trying to do.

Several other lines of work get partway there under different names. One approach drops the idea of a single 'taste' vector per user and models several personas, weighting them by which one a candidate item serves Can attention mechanisms reveal which user taste explains each recommendation?. That is a step toward recognizing that you're doing different things at different times, but personas describe who you are, not where you're headed. Another approach lets you state a preference in plain language and steer the recommender at the moment you use it Can users steer recommendations with natural language at inference?. The goal then enters as an input instead of being learned from history. In practice, if the system can't infer your journey, you can tell it.

Conversational recommenders come at the problem from another angle. They treat a session as a path to optimize rather than a series of separate guesses, learning when to ask a question and when to recommend Can unified policy learning improve conversational recommender systems?. Related work shows that a handful of well-chosen questions can pin down a person's preferences quickly Can user preferences be learned from just ten questions?. Both treat understanding a goal as something you actively uncover, not something you passively predict.

There's a catch in how these systems are trained. Even when language models are trained with reinforcement learning, the rewards are usually standard ranking scores like NDCG and Recall Can recommendation metrics train language models directly?. These scores measure whether you engaged with the right item next. The models learn effective search behavior from this feedback Can LLMs recommend products without ever seeing the catalog?, but the target is still next-item accuracy. Unified text-based recommenders that handle many tasks at once Can one text encoder unify all recommendation tasks? have the same limit: more flexible, same objective.

The surprise is that the obstacle isn't model capacity. The corpus suggests that design choices matter more than model size What architectural choices actually improve recommender system performance?, and journeys are a design choice nobody has made the training target. Language models can already describe what you're working toward. The open question is whether anyone will reward a recommender for helping you get there. The corpus has only one note that studies journeys directly, so treat this as an emerging thread rather than a settled finding.


Sources 9 notes

Can language models discover what users actually want from activity logs?

66% of users pursue valued interest journeys lasting over a month, described in specific phrases like 'designing hydroponic systems for small spaces.' LLM-powered journey discovery bridges the semantic gap that collaborative filtering cannot reach, operating at user-level granularity with persona-level precision.

Can attention mechanisms reveal which user taste explains each recommendation?

AMP-CF represents each user as multiple latent personas weighted dynamically by candidate item. This makes recommendations both diverse and interpretable—each suggestion traces to the specific persona preference it satisfies—without requiring post-hoc reranking.

Can users steer recommendations with natural language at inference?

Mender conditions sequential recommenders on natural-language preferences extracted from reviews, enabling users to steer recommendations at inference without fine-tuning. This approach succeeds on preference-following tasks where traditional recommenders fail because preferences are runtime inputs, not training targets.

Can unified policy learning improve conversational recommender systems?

Research shows that formulating attribute-asking, item-recommending, and timing decisions as a single graph-based RL policy achieves better joint optimization than isolated components. Separation prevents gradient signals from informing one another and fails to optimize conversation trajectory holistically.

Can user preferences be learned from just ten questions?

PReF learns base reward functions from preference data, then uses active learning to select maximally informative questions that reduce coefficient uncertainty. Users can be personalized via inference-time reward alignment without weight modification.

Show all 9 sources
Can recommendation metrics train language models directly?

Rec-R1 demonstrates that LLMs can be trained directly on rule-based recommendation metrics like NDCG and Recall as RL reward signals, eliminating the need for SFT distillation from proprietary models while remaining model-agnostic across different retriever architectures.

Can LLMs recommend products without ever seeing the catalog?

Rec-R1 experiments show that LLMs trained via RL with recommender metrics as rewards can generate effective product search queries without catalog access. The model learns query refinement indirectly through system feedback, paralleling how humans search without knowing platform inventory.

Can one text encoder unify all recommendation tasks?

P5 converts user-item interactions and metadata into natural language and trains a single encoder-decoder across five recommendation task families, matching task-specific models while achieving zero-shot transfer to new items and domains. Unification trades efficiency for composability.

What architectural choices actually improve recommender system performance?

Research shows that architectural choices like removing hidden layers, enforcing constraints on self-similarity, and using appropriate likelihood functions deliver better results than deeper or more complex models. This suggests that problem-specific design decisions matter more than raw representational capacity.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.