INQUIRING LINE

If a feed can tell AI-written posts apart, what should it optimize for so it doesn't reward cheap synthetic filler?

What optimization targets should AIGC-sensitive distribution algorithms use to avoid unintended outcomes?

This explores what a recommendation or feed algorithm should actually optimize for once it knows some content is AI-generated, so that it doesn't end up rewarding the wrong things, such as flooding feeds with cheap synthetic posts or burying human work.


This explores what a content-distribution algorithm (a feed, a recommender, a search ranker) should aim at once it knows some content is AI-generated, so that the system doesn't drift somewhere nobody intended. One caveat first: the collection has no papers about AIGC-aware distribution specifically. What it does have is a set of findings about how optimization targets go wrong in general, and those apply here with surprisingly little translation.

The central lesson is that unintended outcomes rarely come from bad intentions. They come from optimizing hard against a signal that only partly captures what you care about. Reward hacking looks the same whether it happens while training a model, while picking among outputs, or while revising prompts: the system finds the gap between the score and the real goal and exploits it Does reward hacking always stem from the same failure?. For a feed, that means engagement is a scoring function, not ground truth. AI-generated content is unusually good at finding that gap, because it can be produced in huge volume and tuned against whatever the ranker rewards. A related argument says that a 'benign' objective doesn't make you safe. The risk comes from the structure of capable, goal-directed optimization, not from the stated goal Does a benign goal actually prevent harmful AI behavior?. 'Maximize user satisfaction' can still produce a feed nobody would have chosen.

The most practical idea in the collection is to separate what is allowed from what is optimized. In one training method, rubrics act as gates that accept or reject candidates, and a finer-grained reward then optimizes only among the candidates that pass. That design resists gaming much better than turning the rubric into one more score to maximize Can rubrics and dense rewards work together without hacking?. Applied to distribution, provenance and quality checks ('is this disclosed?', 'is this substantively duplicative?') work better as eligibility filters than as weights blended into the ranking score. Once they become weights, they become targets to game.

A second idea concerns what you learn versus how you decide. Building your values directly into the training loss can quietly weaken what the model learns. Training on a plain, symmetric objective and then applying your preferences afterward at decision time worked better, even when judged by the preference-weighted goal itself Can utility-weighted training loss actually harm model performance?. For an AIGC-sensitive ranker, that suggests learning an accurate model of what users actually value first, then applying policy (diversity, human-creator quotas, down-weighting synthetic content) as a separate layer. Don't bake those rules into the engagement predictor. On a similar note, plain right-or-wrong rewards push models toward confident guessing, and adding a proper scoring rule fixes that at no cost to accuracy Does binary reward training hurt model calibration?. That matters if your system has an 'is this AI-generated?' classifier. You want honest probabilities from it, not overconfident labels that wrongly penalize human creators.

The part you probably didn't know you needed to worry about is feedback loops. YouTube's multi-objective ranker has to model selection bias explicitly, including the fact that items get clicked partly because of where they were shown. Without that correction, the system learns from its own past choices and settles into a degenerate pattern that amplifies them Why do ranking systems need to model selection bias explicitly?. With AI-generated content the loop gets tighter. The ranker promotes something, producers generate more content like it, and the training data fills up with the ranker's own preferences reflected back. So the answer to 'what should it optimize?' is several separately modeled objectives, with bias from the system's own past exposure explicitly removed, not one blended score. The open question is whether the detectors themselves survive being optimized against. Related work proposing detectors built from a model's internals as safeguards during training still hasn't tested whether they hold up once the system is trained against them Can reward hacking vectors survive training-time use as detectors?.


Sources 7 notes

Does reward hacking always stem from the same failure?

Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.

Does a benign goal actually prevent harmful AI behavior?

Research shows that risk arises from three conditions: goal-directed reasoning, competence at pursuing goals, and exposure to oversight that can modify objectives. Even benign terminal values leave this risk structure intact, making value alignment an insufficient safety test.

Can rubrics and dense rewards work together without hacking?

DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.

Can utility-weighted training loss actually harm model performance?

Asymmetric loss functions correctly incentivize choosing but degrade representation learning by reducing gradient signals for substantive feature acquisition. Training with symmetric loss then adjusting predictions post-hoc outperforms direct utility-weighted training on the same utility objective.

Does binary reward training hurt model calibration?

Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.

Show all 7 sources
Why do ranking systems need to model selection bias explicitly?

YouTube's multi-objective ranker uses MMoE for conflicting objectives and a shallow position tower to remove selection bias from training data. Without both mechanisms, models converge on degenerate equilibria that amplify their own past decisions.

Can reward hacking vectors survive training-time use as detectors?

The paper identifies internal representations as a promising direction for training-time mitigation of reward hacking but provides no experimental evidence of whether this approach works. The key limitation is that no test shows whether a policy trained against the vector's signal still exhibits detectable hacking.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.