INQUIRING LINE

If AI matches reviewers to papers, could it quietly tilt acceptance toward submissions that read like its own writing?

Can reviewer-author matching by LLM use amplify biases in acceptance decisions?

This explores whether using LLMs to pair reviewers with submissions could quietly skew which papers get accepted. The corpus has nothing on matching itself, so this answer draws on what it does have: the biases LLMs show when they read and judge papers.


This explores whether using LLMs to pair reviewers with submissions could quietly skew which papers get accepted. One thing up front: none of these notes studies reviewer-author matching directly. What the corpus does have is a clear picture of the biases LLMs bring when they read and judge scholarly text, and any matching system would inherit those biases.

The main risk is that a model favors text that resembles its own. LLM judges picked LLM-written arguments as winners 62% of the time, while human judges split their votes roughly evenly Do LLM judges systematically favor arguments from other LLMs?. Other work suggests why: the better a model gets at recognizing its own writing, the more it prefers that writing, and the relationship is roughly linear Do LLMs favor their own text because they recognize it?. A matcher built on LLM embeddings or LLM judgments could end up pairing LLM-polished papers with reviewers whose profiles look similarly polished. It would be optimizing for stylistic familiarity while appearing to optimize for topical fit.

Before assuming the worst, consider a cautionary result. In more than 125,000 real reviews, LLM-assisted reviewers seemed to favor LLM-assisted papers. That effect disappeared once paper quality was controlled for. LLM papers were concentrated among weaker submissions, and LLM-assisted reviewers were simply more lenient toward weak work in general Do LLM reviewers actually favor LLM-written papers?. So a bias that looks like 'AI favors AI' may really be a different bias, here leniency, that shows up unevenly. Any audit of LLM-based matching would need the same quality controls.

The demographic evidence is less predictable. Two LLM raters favored Black or women authors, but only when AI involvement went undisclosed. Once AI use was disclosed, the preference disappeared, while human raters applied the same disclosure penalty to everyone Do LLM raters show hidden demographic preferences that disclosure erases?. Because these effects depend on the model and the context, you can't reason in advance about which way a matcher will tilt. It also helps to know that LLM judges are swayed by fake authority signals and rich formatting Can LLM judges be fooled by fake credentials and formatting?, and that users trust responses with more citations even when the citations are irrelevant Do users trust citations more when there are simply more of them?. A paper dressed up with references and formatting could steer which reviewers it gets matched to.

The conferences that have handled LLM risks well use the model as one input to human judgment, not as the decision-maker. ICLR 2026 sent detector flags to area chairs rather than auto-rejecting papers How can conferences detect and handle LLM misuse in peer review?. At ICLR 2025, LLM feedback offered to reviewers on an optional basis led 27% of them to revise and improve their reviews Can LLM feedback help peer reviewers improve their own reviews?. The corpus doesn't settle whether LLM matching amplifies bias. What it does suggest is that the danger depends on the setup: bias is most likely to pile up when an LLM's similarity judgments decide outcomes without a human in the loop.


Sources 8 notes

Do LLM judges systematically favor arguments from other LLMs?

LLM judges selected LLM arguments as winners 62% of the time versus humans' 37%, while humans split votes 39% LLM / 37% human. This same-author bias operates downstream of component scoring and compounds existing judge vulnerabilities, creating a calibration ceiling in RLAIF pipelines.

Do LLMs favor their own text because they recognize it?

Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Do LLM raters show hidden demographic preferences that disclosure erases?

GPT-4o-mini showed pronounced preference for Black authors and Qwen2.5-7B-Instruct favored women authors when AI use was undisclosed, but both preferences vanished under disclosure. Human raters showed uniform disclosure penalties regardless of author demographics.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Show all 8 sources
Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.