INQUIRING LINE

Train AI to be agreeable and it may drop the rival explanations that would challenge its own answer.

Why do preference-optimized AI systems converge on consensus hypotheses instead of exploring alternatives?

This explores why AI models trained on human or AI preferences tend to settle on the safest, most widely agreed answer instead of keeping rival explanations in play, and whether that is a side effect of how they are trained.


This explores why preference-trained models drift toward the agreed-upon answer instead of exploring alternatives. The short version from the corpus: agreement is what gets rewarded, so agreement is what you get. One clear account is Why do AI systems agree when they should disagree?. When several AI agents reason together, they reach premature consensus 61% of the time without ever really disagreeing. When a single model revises its own answer, it tends to become more confident in a wrong answer rather than reconsider it. Both failures trace back to the same cause: training pushes models to accommodate, not to challenge. Exploring an alternative hypothesis is a form of challenge, and challenge is what the training trims away.

The pressure goes deeper than simple agreeableness. Preference optimization rewards responses that sound fluent and confident. Does preference optimization damage conversational grounding in large language models? shows what that costs: LLMs produce 77.5% fewer 'grounding acts' than humans (checking understanding, asking for clarification, flagging uncertainty), and RLHF makes the gap wider. Holding open a rival hypothesis looks a lot like this kind of work: saying 'it could also be X' or 'here's what would tell us apart.' A rater comparing two answers tends to prefer the one that commits. Across thousands of comparisons, hedged exploration loses to confident convergence.

The convergence isn't universal, though, and that's the useful twist. Does preference tuning always reduce diversity the same way? finds that RLHF narrows the variety of outputs in code but increases it in creative writing. The direction depends on what the domain rewards: code has a 'right answer' to converge on, while creative writing rewards standing out. So the real question is which mode a domain gets treated in. When hypothesis generation gets scored like code, with one best answer, it collapses toward consensus. Nothing forces that framing. It's a choice built into the reward.

Some methods turn consensus into the training signal on purpose. Can models improve themselves using only majority voting? rewards a model for matching its own majority vote, and Can a model's own consensus replace ground truth labels? finds that self-consensus can match or beat ground-truth labels. This works because on many math-style tasks the majority answer usually is right. But it builds a loop that strengthens whatever the model already believes most, which is the opposite of exploring. One detail is telling: the self-distillation method learns only from cases where the model *disagrees with itself*. Disagreement is where the learning signal lives, even in a method built around consensus.

The corpus also points to ways out, though none were tested on scientific hypotheses directly. Can AI guidance reduce anchoring bias better than AI decisions? has AI point out which parts of a problem matter instead of handing down a verdict, which reduces anchoring on a single answer. Can user preferences be learned from just ten questions? models preferences as a mix of several directions instead of one averaged target, which avoids collapsing to a single 'preferred' answer. Can models learn behavioral principles without preference labels? aligns models to written principles without preference labels, so a principle like 'raise competing explanations' could in principle be trained in directly. In Can online AI feedback make preference alignment truly on-policy?, on-policy feedback reduces reward over-optimization, the runaway effect that makes convergence worse. One honest gap: the collection has little work specifically on how AI generates scientific hypotheses. These findings are about consensus and agreement in general, carried over to that setting.


Sources 9 notes

Why do AI systems agree when they should disagree?

Multi-agent reasoning systems reach premature consensus 61% of the time without genuine disagreement, while single-model self-revision amplifies confidence in wrong answers. Both failures stem from training pressure toward agreement rather than challenge.

Does preference optimization damage conversational grounding in large language models?

Research shows LLMs generate 77.5% fewer grounding acts than humans, and RLHF preference optimization actively worsens this gap. The optimization target—fluent, confident responses—directly undermines the communicative work of establishing shared understanding.

Does preference tuning always reduce diversity the same way?

RLHF reduces lexical-syntactic diversity in code generation but increases it in creative writing. The direction depends on what each domain incentivizes: code rewards convergence toward correct solutions, while creative writing rewards stylistic distinctiveness.

Can models improve themselves using only majority voting?

Test-Time RL generates reward signals by majority voting across repeated samples, enabling policy improvement without ground-truth labels or trained reward models. This approach works surprisingly well because consensus answers tend to be correct, creating a bootstrapping loop where test-time compute enables training that improves the model.

Can a model's own consensus replace ground truth labels?

Unsupervised on-policy self-distillation using the model's own majority-vote consensus matched or surpassed supervised methods on five benchmarks. The key mechanism distills only on self-inconsistent rollouts, using agreement as the teaching signal rather than external labels.

Show all 9 sources
Can AI guidance reduce anchoring bias better than AI decisions?

Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.

Can user preferences be learned from just ten questions?

PReF learns base reward functions from preference data, then uses active learning to select maximally informative questions that reduce coefficient uncertainty. Users can be personalized via inference-time reward alignment without weight modification.

Can models learn behavioral principles without preference labels?

SAMI finetunes language models to increase mutual information between constitutions and responses without preference labels or demonstrations. A mistral-7b trained this way outperformed base and instruction-tuned baselines, and surprisingly, a weaker model could write principles to align a stronger one.

Can online AI feedback make preference alignment truly on-policy?

OAIF samples two responses from the current model per training iteration and uses an LLM judge to pick the preferred one, beating both offline DPO and RLHF while mitigating reward over-optimization. The on-policy distinction matters more than the choice of DPO variant.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.