Do LLM reviewers favor papers written by other LLMs?
When LLMs serve as peer reviewers, do they systematically score papers differently based on whether they were written by humans or LLMs? Understanding this matters for fair and trustworthy publication systems.
The LLM-REVal paper builds a multi-round simulation of publication. A research agent writes papers, human-authored and LLM-authored alike, and a review agent scores them through a five-stage pipeline built on AgentReview. The central finding is that "LLM reviewers systematically inflate scores for LLM-authored papers, assigning them markedly higher scores than human-authored ones," and that they "persistently underrate human-authored papers with critical statements (e.g., risk, fairness), even after multiple revisions." Human annotation of the same outputs points the other way: "Human reviewers do not prefer LLM papers over human papers as LLM reviewers do, and those consistently low-scored human papers are deemed valuable."
The authors attribute the pattern to two biases: "a linguistic feature bias favoring LLM-generated writing styles, which is more concise, lexically diverse, and complex, and an aversion toward critical statements." The discussion adds a framing effect. Among human-authored papers, more negative keywords go with more positive abstract sentiment and lower review scores. Among LLM-authored papers, sentiment stays positive regardless, and scores rise with the count of such keywords. The authors read this as "a negative framing of critical topics could exacerbate bias in LLM reviews, whereas a positive framing of similar topics tended to yield disproportionately higher scores." The same excerpt also reports a counterweight: revisions guided by LLM reviews gained quality in both LLM-based and human evaluations.
Against the nearest notes, this excerpt extends How much does rhetorical style shift AI review scores?. Both show that LLM scores track presentation. This excerpt places part of that sensitivity in how the reviewer handles critical content and in its taste for LLM-style prose, not in evidence framing or novelty stance, which it does not test. The taste for LLM-style prose parallels Can readers tell LLM abstracts from human ones?, where human readers also rated LLM-edited text well, though in a survey experiment rather than a simulated reviewer. The contrast with Can structured pipelines make LLM novelty assessment reliable? matters. That pipeline's authors report strong agreement with human reviewers on novelty. This excerpt finds misalignment on style and critical framing, so agreement on one judgment does not transfer to others.
The excerpt does not establish how often deployed LLM reviewers behave this way. The agents are simulated, and the research agent predicts experimental results rather than running experiments, so the papers are not executed science. The critical-topic finding rests on keyword detection and abstract sentiment, and it reports correlations, not controlled tests. The excerpt also omits annotation sample sizes and agreement statistics, so the human-alignment claim cannot be weighed from it alone. What the evidence supports is narrower: LLM review scores can depend on writing style and on how critical content is framed, so an LLM score should not be read as an impartial judgment. The authors' own conclusion, that LLM reviewers "cannot yet be fully trusted as impartial evaluators," is the claim the evidence carries most firmly.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do restrictions on reviewer LLM use actually shape peer review behavior?- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Can rules against undisclosed LLM use change reviewer behavior without enforcement?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
same finding that LLM scores track presentation; this excerpt adds critical-content framing and style.
-
Can readers tell LLM abstracts from human ones?
Do readers with ML expertise reliably distinguish human-written, LLM-generated, and LLM-edited research abstracts? Understanding this matters for evaluating whether readers can serve as effective gatekeepers against LLM content.
parallel preference for LLM-style prose, from human readers rather than a simulated reviewer.
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
same family of surface-feature bias, shown here without adversarial prompts.
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
substantial noncompliance makes the simulated scenario plausible; the two studies measure different outcomes.
-
Can structured pipelines make LLM novelty assessment reliable?
Explores whether breaking novelty assessment into extraction, retrieval, and comparison stages helps LLMs align with human peer reviewers and produce more rigorous, evidence-based evaluations.
contrast: reported agreement on novelty does not transfer to style or critical framing.
-
Do LLM reviewers actually favor LLM-written papers?
Does the apparent bias of LLM-assisted peer reviewers toward LLM-generated papers reflect genuine preferential treatment or an artifact of quality distribution? The answer shapes how we interpret reviewer behavior.
contradicts: once paper quality is controlled, the LLM-paper favoritism vanishes, traced instead to LLM papers clustering among weak submissions
-
Can AI systems safely replace human peer reviewers?
Explores whether AI reviewers meet two critical conditions for automation: maintaining diverse perspectives and resisting score manipulation. Tests whether current systems are ready to handle peer review at scale.
evidence for: zero-shot rewrites raise AI reviewer scores without added scientific substance, consistent with style-driven inflation
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Stop Automating Peer Review Without Rigorous Evaluation
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
Original note title
simulated LLM reviewers inflate scores for LLM-authored papers and underrate human papers with critical statements — traced to style and aversion to criticism