INQUIRING LINE

Could AI agents be trained to stop favoring certain brands just by balancing which sources appear most in training data?

Can balancing training data by source eliminate agent source preference bias?

This explores whether you could fix AI agents' habit of favoring certain brands, websites or vendors by balancing the training data so that no source looks better than another, and what the corpus says about where that habit comes from.


This explores whether you could stop AI agents from favoring certain sources (a brand, a retailer, a publisher) by balancing the training data so no source has an unfair edge. The corpus doesn't test balancing directly, but it does show where the bias comes from, which explains why balancing is a reasonable idea and also why it probably isn't enough. Across 12 models and three domains, agents picked an item from a favored source over a better item from another source about two-thirds of the time Do language models favor sources regardless of item quality?. Two findings from that work matter here. When training data pairs a source with good outcomes, the preference appears. When the source label is hidden, the bias weakens but doesn't disappear. So the source name works as a learned shortcut, and part of the bias seems to leak through cues other than the label itself.

This is why balancing by source is a partial fix at best. Balancing the counts doesn't remove the correlation between a source and its outcomes. A source can show up equally often and still be linked with better results, more confident wording or more polished descriptions. A wider critique of 'theory-free' AI makes the same point: models trained only to be accurate absorb whatever correlations the data contains and treat them as if they mean something Can AI models be truly free from human bias?. Source preference is a small, measurable case of that. 'This came from X' stands in for 'this is good.' Balancing the data trims one path to that shortcut. It doesn't teach the agent to check the item against the requirements.

There's also a reason the bias might come back after training. Agents trained on curated demonstrations only learn the habits the curators showed them Can agents learn beyond what their training data shows?. Post-training tends to sharpen a model's existing tendencies rather than widen what it considers Do base models find more solutions than post-trained ones?. If a slight lean toward a source survives balancing, a later fine-tuning round can make it stronger rather than weaker.

The more promising direction in the corpus is to change how the agent decides rather than what it was trained on. Evaluators that gather concrete evidence before judging were about 100 times more consistent than LLM judges that answer in one shot Can agents evaluate AI outputs more reliably than language models?. Applied to source bias, that suggests making agents check each item against the user's requirements one by one before choosing, combined with hiding source labels at decision time. The surprising takeaway is that the bias isn't just a quirk of skewed data. It's what happens when a model can grab an easy cue instead of doing the comparison, so the lasting fix is to make the comparison unavoidable. The corpus doesn't yet show anyone combining these approaches, so that remains an open question.


Sources 5 notes

Do language models favor sources regardless of item quality?

Across 12 models and three domains, agents select items from favored sources even when they satisfy fewer requirements than alternatives. Hiding source labels weakens the bias, and training data that pairs sources with better outcomes induces the preference.

Can AI models be truly free from human bias?

Research shows that 'theory-free' AI models mask bigotry behind high accuracy metrics while committing fundamental statistical errors. A 95% accurate criminal justice system would wrongly convict thousands, demonstrating that model sophistication does not validate causal inference.

Can agents learn beyond what their training data shows?

Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.

Do base models find more solutions than post-trained ones?

Across 14 model pairs and three agentic benchmarks, base models equipped with only relaxed system prompts eventually surpass post-trained counterparts in pass@K coverage as rollout budget grows. Post-training bimodalizes task outcomes, sharpening performance on easy cases while eliminating rare-but-reachable solutions.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.