SYNTHESIS NOTE
Topics›Domain Specialization›this note

Do LLM reviewers favor papers written by other LLMs?

When LLMs serve as peer reviewers, do they systematically score papers differently based on whether they were written by humans or LLMs? Understanding this matters for fair and trustworthy publication systems.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The LLM-REVal paper builds a multi-round simulation of publication. A research agent writes papers, human-authored and LLM-authored alike, and a review agent scores them through a five-stage pipeline built on AgentReview. The central finding is that "LLM reviewers systematically inflate scores for LLM-authored papers, assigning them markedly higher scores than human-authored ones," and that they "persistently underrate human-authored papers with critical statements (e.g., risk, fairness), even after multiple revisions." Human annotation of the same outputs points the other way: "Human reviewers do not prefer LLM papers over human papers as LLM reviewers do, and those consistently low-scored human papers are deemed valuable."

The authors attribute the pattern to two biases: "a linguistic feature bias favoring LLM-generated writing styles, which is more concise, lexically diverse, and complex, and an aversion toward critical statements." The discussion adds a framing effect. Among human-authored papers, more negative keywords go with more positive abstract sentiment and lower review scores. Among LLM-authored papers, sentiment stays positive regardless, and scores rise with the count of such keywords. The authors read this as "a negative framing of critical topics could exacerbate bias in LLM reviews, whereas a positive framing of similar topics tended to yield disproportionately higher scores." The same excerpt also reports a counterweight: revisions guided by LLM reviews gained quality in both LLM-based and human evaluations.

Against the nearest notes, this excerpt extends How much does rhetorical style shift AI review scores?. Both show that LLM scores track presentation. This excerpt places part of that sensitivity in how the reviewer handles critical content and in its taste for LLM-style prose, not in evidence framing or novelty stance, which it does not test. The taste for LLM-style prose parallels Can readers tell LLM abstracts from human ones?, where human readers also rated LLM-edited text well, though in a survey experiment rather than a simulated reviewer. The contrast with Can structured pipelines make LLM novelty assessment reliable? matters. That pipeline's authors report strong agreement with human reviewers on novelty. This excerpt finds misalignment on style and critical framing, so agreement on one judgment does not transfer to others.

The excerpt does not establish how often deployed LLM reviewers behave this way. The agents are simulated, and the research agent predicts experimental results rather than running experiments, so the papers are not executed science. The critical-topic finding rests on keyword detection and abstract sentiment, and it reports correlations, not controlled tests. The excerpt also omits annotation sample sizes and agreement statistics, so the human-alignment claim cannot be weighed from it alone. What the evidence supports is narrower: LLM review scores can depend on writing style and on how critical content is framed, so an LLM score should not be read as an impartial judgment. The authors' own conclusion, that LLM reviewers "cannot yet be fully trusted as impartial evaluators," is the claim the evidence carries most firmly.

Inquiring lines that read this note 6

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do restrictions on reviewer LLM use actually shape peer review behavior?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 103 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

simulated LLM reviewers inflate scores for LLM-authored papers and underrate human papers with critical statements — traced to style and aversion to criticism