Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

Paper · arXiv 2610.03195 · Published October 2, 2026
LLM Agents

ABSTRACT As LLM agents decide on users’ behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources are selected. We study source preference in end-to-end search with 12 agent models across three domains. Comparing items from different sources that satisfy the same requirements at the same position, we find that each model prefers some sources and avoids others in every domain, largely agreeing on which. This preference can outweigh how well items satisfy the request: an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better one from a dispreferred source, but almost never in the reverse case. The information identifying an item’s source affects selection by itself: hiding it weakens the preference, and relabeling an item with a preferred source raises its selection rate. We test two routes to this preference: training that rewards better items can make a source a shortcut for requirement satisfaction, and missing information can trigger preconceptions about the source. Supplying missing information or a prompt countering these preconceptions reduces source preference.

Introduction. Large language model agents increasingly move beyond providing information to deciding what users see. They decide which products to show (Yao et al., 2022a; Zhou et al., 2024), which accommodations to recommend (Xie et al., 2024; Hadad et al., 2026), and which papers to present (Skarlinski et al., 2024; Asai et al., 2026; Shen et al., 2026). Consider an agent searching for accommodations from sources such as Booking.com and Expedia. The agent has a source preference when it favors items from one source over another, even though they satisfy the request equally well. Such a preference can narrow the options available to users and lead agents to pass over better items because of their source. If these selections subsequently feed back into training, existing preferences could become stronger, concentrating visibility on favored sources while leaving others overlooked even when they offer equally good items. Prior work provides evidence of such preferences, but has not examined this influence in real-world settings where agents retrieve and select items themselves (Dai et al., 2025; Khan et al., 2026; Schuster et al., 2026). For example, Khan et al. (2026) present a few items with semantically identical content attached to different sources and show that models favor some sources over others. In actual search, however, each source offers its own items, which differ in content and how well they satisfy the user’s request. We therefore study whether source preference appears in end-to-end search, whether the information identifying an item’s source itself contributes to such preferences, and how this reliance on the source might arise and be reduced. Our study covers 12 agent models and three domains: shopping, accommodation, and scholarly search.

Establishing source preference in end-to-end search requires accounting for differences in the retrieved items. A higher selection rate alone is insufficient: a source’s items may better satisfy the request or appear in more favorable positions (Zheng et al., 2023; Allouah et al., 2025). We therefore compare items from different sources that satisfy the same requirements, with position held fixed, to score each source’s preference and classify it as preferred or dispreferred (§3). Under these controls, every model prefers some sources and avoids others in every domain, with most frequently retrieved sources falling into either group. Models largely agree on these preferences: most models prefer Booking.com, while many avoid Expedia, even though the items from these sites satisfy the same requirements. We next examine whether these preferences persist when the preferred source offers a less satisfying item. To do so, we compare two items, one of which satisfies one fewer requirement than the other. Agents select the less satisfying item about two-thirds of the time when it comes from a preferred source and the better item from a dispreferred source. They almost never select the less satisfying item when it comes from a dispreferred source and the better item from a preferred source. This suggests that source preference can leave users with less satisfying items and cause dispreferred sources to lose selections even when they offer better ones (§4).

These comparisons establish source preference under the same conditions, but do not isolate the contribution of the information identifying an item’s source. We therefore examine whether changing this information changes selection while keeping the items fixed. First, we present the same search result lists with this information hidden and then restored: hiding it weakens source preference, and restoring it widens the gap between preferred and dispreferred sources. We then relabel the same items with preferred and dispreferred sources, keeping their titles and content fixed: across every model and domain, selection rates are higher when the same content is labeled with a preferred source than with a dispreferred one. These interventions show that this information itself contributes to agents’ selections (§5).

Given that the source itself affects selection, why might agents use it as a selection signal? We examine how this reliance can arise during training and be triggered at inference. During training, an agent may learn to use the source as a shortcut for requirement satisfaction: when a source is more often paired with the better item during DPO (Rafailov et al., 2024), agents learn to prefer it, while balancing this pairing can mitigate an existing preference (§6). At inference, missing information may trigger the agent’s preconceptions about the source, which fill the gap; supplying the missing information or prompting the agent to counter such preconceptions reduces this reliance (§7).

Method. 2 END-TO-END AGENT SETTING Agents and Search Workflow. We evaluate GPT-{5.4-nano (OpenAI, 2026a), 5.6-Luna (OpenAI, 2026b)}, Gemini-3.7-Flash (Deepmind, 2026), GLM-5.3-Flash (GLM-5-Team et al., 2026), DeepSeek-v4-Flash-0731 (DeepSeek-AI et al., 2026), Llama-4-{Maverick, Scout} (Meta, 2025), Llama-3.1-8B-Instruct (Grattafiori et al., 2024), Tulu-3-8B (Lambert et al., 2025), Qwen3.5- 27B (Qwen, 2026), Qwen3-30B-A3B (Yang et al., 2025), and Qwen2.5-32B-Instruct (Qwen et al., 2025) as agents in an end-to-end search-and-selection setting. All agents follow a ReAct-style interaction loop (Yao et al., 2022b), alternating reasoning and actions with observations from the search environment under a shared prompt. Given a user request, the agent issues a search query and receives up to 10 results, each containing a title, content, and URL. The agent then evaluates these results against the request and may select multiple items or issue another query if none are suitable. We define each item’s source as its URL’s registrable domain and measure source preference by comparing source exposure and selection in the trajectories (Implementation details are in §B).

Evaluation Domains. We evaluate agents across three domains where agents commonly make selections: Shopping, Accommodation, and Scholar. To obtain realistic requests with specific requirements, we use WebShop (Yao et al., 2022a) for shopping products, HotelQuEST (Hadad et al., 2026) for accommodation searches, and ScholarGym (Shen et al., 2026) for searching scholarly works. We use 4,822 requests: 1,500 from WebShop, 786 from HotelQuEST, and 2,536 from Schol- MEASURING SOURCE PREFERENCE Running example: is a amazon.com preferred?

1 Request Shopping WebShop Accommodation HotelQuEST Scholar ScholarGym e.g.

“I am looking for earphones with a mic in the color blue, with a price lower than $40.”

2 Search and selection blue earphones with mic under $40 Search results title content URL Blue Earphones with Mic... Built-in mic, in-ear earphon... a amazon.com/s?k=blue+... selected 2 walmart.com w Earphones with Mic... not selected 3 ebay.com e Sport Earphones, Blue... selected 3 Requirement satisfaction source-blind Judge input title + content snippet (URL, source names removed) Satisfied requirements q earphones mic blue < $40 a Blue Earphones... w Earphones with... e Sport Earphones... requirement-matched pair: same satisfied requirements 4 Matched comparisons Compare the requirement-matched pair a vs w (from 3 ) at the same position 1. Rotate so items visit every position Run 1 Run 2 Run 3 Run 1 = the list in 2 2. One comparison per position p a vs w D 1 +1 2 0 3 −1 D: +1 only a selected 0 both or neither −1 only w selected 5 Preference score a ’s D, by opponent vs w (from 4 ) +1 0 −1 vs other sources +1 +1 0 ... adjust for who a faced: Bradley–Terry model rating βa (0 = average source) τa = [D | βa, βopp = 0] = +0.12 a selected +12 pp vs. an average source 6 Source classification Is βa significantly above 0? (request-level bootstrap, FDR 5%) SOURCE+ (preferred) Rule for every source β significantly > 0 SOURCE+ (preferred) no clear evidence SOURCE0 (neutral) β significantly < 0 SOURCE−(dispreferred) Requirement Satisfaction. We assess requirement satisfaction for every search result exposed to the agent using a source-blind judge (§3; §4.2). For Shopping, we use WebShop’s structured requirements; for Accommodation and Scholar, an LLM extracts verifiable requirements from each request. An LLM judge evaluates each requirement using only the result’s title and content snippet, with URLs and source names removed from both. For validation, human annotators provide requirement-level labels for a sample of agent-selected results in each domain. Across domains, judge–human agreement is comparable to inter-annotator agreement (Krippendorff’s α: 0.65–0.75 vs. 0.67–0.72). §D provides methodological details and per-domain validation results.

Conclusion. We studied source preference in end-to-end search across 12 agent models and three domains. Agents favor sources even when items satisfy the same requirements, and these preferences can accompany selections of less satisfying items. Controlled experiments establish that source identity itself affects selection, while preference training shows how source–satisfaction associations can create or reduce source preference. Supplying missing information and prompting agents to reconsider source-based assumptions reduce preference. These findings highlight the need to evaluate agents not only by which items they select, but also by how source identity shapes those selections.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do training data quality and composition affect downstream model performance? How do agents learn to distinguish valuable feedback from noise? How can we detect and account for LLM involvement in academic writing? How does AI-generated content create social proof without authentic interaction? When do simpler collaborative filtering approaches outperform complex LLM recommenders? How reliably can language models perform causal versus temporal reasoning? How should recommendation systems balance individual preference and diversity? Why do training associations persist despite contradictory contextual information? How do clinicians calibrate trust in AI medical recommendations? How can we reduce inherent biases in LLM-based evaluation judges? How can AI systems reliably guide voters without introducing political bias? Does preference optimization undermine conversational grounding in language models?