LLM-REVal: Can We Trust LLM Reviewers Yet?

Paper · arXiv 2510.12367 · Published October 14, 2025
Domain Specialization in LLMs

The rapid advancement of large language models (LLMs) has inspired researchers to integrate them extensively into the academic workflow, potentially reshaping how research is practiced and reviewed. While previous studies highlight the potential of LLMs in supporting research and peer review, their dual roles in the academic workflow and the complex interplay between research and review bring new risks that remain largely underexplored. In this study, we focus on how the deep integration of LLMs into both peer-review and research processes may influence scholarly fairness, examining the potential risks of using LLMs as reviewers by simulation. This simulation incorporates a research agent, which generates papers and revises, alongside a review agent, which assesses the submissions. Based on the simulation results, we conduct human annotations and identify pronounced misalignment between LLM-based reviews and human judgments: (1) LLM reviewers systematically inflate scores for LLM-authored papers, assigning them markedly higher scores than human-authored ones; (2) LLM reviewers persistently underrate human-authored papers with critical statements (e.g., risk, fairness), even after multiple revisions. Our analysis reveals that these stem from two primary biases in LLM reviewers: a linguistic feature bias favoring LLMgenerated writing styles, and an aversion toward critical statements. These results highlight the risks and equity concerns posed to human authors and academic research if LLMs are deployed in the peer review cycle without adequate caution. On the other hand, revisions guided by LLM reviews yield quality gains in both LLMbased and human evaluations, illustrating the potential of the LLMs-as-reviewers for early-stage researchers and enhancing low-quality papers. The code is available at https://github.com/PlusLabNLP/LLM-REVal.

Introduction. Large language models (LLMs) are demonstrating growing autonomy in scientific research, spanning tasks such as idea generation (Wang et al., 2024; Yang et al., 2024; Qi et al., 2023; Si et al., 2025; Kumar et al., 2024), automated citation (Press et al., 2024; Ajith et al., 2024; Kang & Xiong, 2024), and data analysis (Huang et al., 2024; Tang et al., 2023; Guo et al., 2024; Tian et al., 2024; Chan et al., 2025; Nolte & Tomforde, 2025). With a substantial increase in submissions and unprecedented pressure on peer review (AAAI 2025 received over 23,000 submissions1). Using LLMs to alleviate review workload (Gao et al., 2024b; D’Arcy et al., 2024; Zhu et al., 2025b) has gained interest. Notably, recent work demonstrated that LLMs can enhance the peer review process by improving the clarity, actionability, and interactivity of reviews (Liang et al., 2023; Thakkar et al., 2025), and by generating high-quality meta-review summaries (Hossain et al., 2025).

Despite their potential, the use of LLMs as reviewers raises profound concerns for the integrity of the scholarly ecosystem (Du et al., 2024; Lin et al., 2025; Zhu et al., 2025a), as their judgments may embed implicit biases that compromise evaluative fairness. Existing studies have identified several vulnerabilities of LLM reviewers, including susceptibility to hidden prompt manipulation and a focus on self-disclosed limitations in manuscripts (Ye et al., 2024; Tyser et al., 2024).

In practice, on the one hand, LLMs are increasingly used to refine phrasing or even generate manuscripts. On the other hand, some reviewers, despite explicit prohibitions, delegate their responsibilities to LLMs (Yu et al., 2025), creating the prospect that an LLM-refined or even LLM-written paper may be reviewed by an LLM reviewer. The systemic implications and new risks arising from the dual roles and feedback loop of LLMs in the academic process remain largely unexplored.

To address this, we propose LLM-REVal (LLM REViewer Re-EValuation) through a multi-round simulation of the academic publication process. As shown in Figure 1, the simulation begins with a research-review round that includes both human- and LLM-authored submissions, followed by multiple revise-review rounds where low-scoring papers are resubmitted in revised form. The simulation comprises a Research Agent and a Review Agent. The Research Agent autonomously executes the research workflow, encompassing literature retrieval, idea generation, experimental design, result analysis, manuscript compilation, and subsequent possible revision. The generated papers are visually and structurally indistinguishable from human-authored papers. The Review Agent emulates the complete scholarly peer-review pipeline, including initial reviewer assessments, rebuttals, reviewer reassessments, meta-reviews, and final decisions. The review agent exhibits indicative ability for human paper acceptance and correlates well with human review scores, ensuring the reliability of the simulation.

Focusing on potential risks, particularly systemic biases arising from deep integration of LLMs into scholarly workflows, we quantitatively analyze multi-round review results, comparing LLM-authored and human-authored papers, as well as original low-scoring submissions and their revisions. We identify salient patterns in LLM reviewer behavior, contrast these with human judgments, and trace the underlying sources of bias. Our findings are as follows:

• Through analyzing LLM-reviewer ratings, we identify 1) LLM reviewers assign significantly higher scores to LLM papers compared to human papers on the same topic. 2) Revisions show significant score improvements over their low-scoring initial submissions, regardless of authorship. 3) Despite multiple revisions, the LLM reviewer persistently assigns low scores to certain types human papers.

• The comparisons with human evaluation reveal a misalignment between LLM reviews and human judgments: Human reviewers do not prefer LLM papers over human papers as LLM reviewers do, and those consistently low-scored human papers are deemed valuable. This misalignment indicates systematic biases in LLM reviewers.

• We identify a linguistic feature bias favoring LLM-generated writing styles, which is more concise, lexically diverse, and complex, and an aversion toward critical statements (e.g., discussions about risk, fairness). Consequently, submissions exhibiting LLM-style linguistic features tend to receive inflated scores, whereas work containing critical statements may be severely undervalued.

These findings highlight the risks and fairness concerns posed to human authors with irresponsible deployment of LLMs in peer review, particularly as LLMs become increasingly integrated across multiple stages of the research lifecycle. Without proper safeguards, the adoption of LLMs as reviewers risks entrenching systemic biases and eroding trust in the peer-review process.

Related work. AI-assisted Review The recent surge in manuscript submissions has placed unprecedented pressure on the peer review process, motivating growing interest in leveraging LLMs to alleviate this burden. Prior empirical studies have reported promising results in adopting LLMs as reviewers. Liang et al. (2023) proposes that GPT-4 can provide useful feedback on research papers. Through a randomized study of 20,000 ICLR 2025 reviews, Thakkar et al. (2025) finds that LLM-assisted feedback significantly improves the clarity, actionability, and interactivity of peer reviews. In addition, Hossain et al. (2025) investigates the application of LLMs in meta-reviews, finding that they are excellent at multi-perspective summaries of reviews. Moreover, several studies have started to develop reviewer agents, including static review generation system such as Reviewer2 (Gao et al., 2024b), MARG (D’Arcy et al., 2024), DeepReview (Zhu et al., 2025b) and dynamic review system such as Agentreview (Jin et al., 2024) and ReviewMT (Tan et al., 2024). Despite these advances, LLM reviewers exhibit notable limitations. They often overemphasize limitations explicitly stated in manuscripts and assign inflated scores to submissions with limited substantive content (Ye et al., 2024). Further, they display low sensitivity to ethical concerns and technical nuances (Tyser et al., 2024), and are vulnerable to adversarial text perturbations (Lin et al., 2025) as well as hidden prompt manipulations (Ye et al., 2024). While these studies have revealed specific risks of LLM-based reviewing, they have largely treated the review stage in isolation, neglecting its coupling with other phases of the scholarly lifecycle.

AI-assisted Research The integration of LLMs into scientific research has driven substantial progress at multiple stages of the research workflow, including idea generation (Wang et al., 2024; Yang et al., 2024; Qi et al., 2023; Si et al., 2025; Kumar et al., 2024; Li et al., 2024a), automated citation (Press et al., 2024; Ajith et al., 2024; Kang & Xiong, 2024), experiment design, execution and analysis (Nolte & Tomforde, 2025; Huang et al., 2024; Tang et al., 2023; Tian et al., 2024; Chan et al., 2025; Guo et al., 2024), among others. These studies validate the feasibility and value of LLMs for scientific automation. Si et al. (2025) found that LLM-generated ideas were assessed as superior in novelty compared to those conceived by human experts. LitSearch (Ajith et al., 2024) and CiteME (Press et al., 2024) reveal both the potential and limitations of LLMs in handling complex scholarly information retrieval and precise citation matching.

Method. Our work simulates the academic publication process through iterative interactions between a research agent and a review agent. The research agent generates ideas, conducts studies, and revises manuscripts, while the review agent evaluates submissions and provides reviews. This multi-round framework enables analysis of LLM-mediated review dynamics and associated risks.

Drawing inspiration from Lu et al. (2024) and Si et al. (2025), we constructed a research agent that initiates the target research with a list of keywords, which define the research direction. According to keywords, the research agent conducts literature retrieval, generates ideas, designs experiments, predicts and analyzes results, drafts the paper, and finally compiles the manuscript from LATEX to PDF format.

Literature Retrieval Retrieval-augmented generation (Gao et al., 2024a) with literature is integrated at multiple stages of the research agent workflow to ensure knowledge accuracy. During idea generation, we query the Semantic Scholar API (Kinney et al., 2023) with keywords and rank retrieved papers by relevance, empirical grounding, and novelty. The top-ranked papers are then used to inspire research ideas and guide experimental design. In the paper writing stage, previously retrieved papers are aggregated into a topic-specific paper bank. We filter this corpus by semantic similarity to select references for drafting. Additionally, the Google Search API2 is used to obtain relevant web content, which is summarized to support background sections.

Idea Generation and Experimental Design Based on the retrieved literature and manually curated examples, the research agent generates candidate ideas. Cosine similarity is used to remove duplicate ideas, after which the remaining ideas are ranked, and the top idea is selected. Building on this idea, the research agent develops an experiment plan and refines it according to demonstrations summarized from the retrieved literature. Limited by LLM coding capability and time-consuming experiment execution, which may make the simulation unable to scale, we employ the research agent to predict results rather than obtain results through experiment iteration.

Paper Writing We design an effective and efficient workflow that combines rule-based components with LLM-based generation to guide the paper writing process. 1) Plug-and-Play Template. The research agent is initialized with the official ICLR LATEX template. Following standard academic structure, the paper is organized into: Abstract, Introduction, Background, Method, Experimental Setup, Results Analysis, Related Work, and Conclusion. 2) Sequential Iteration. Sections are generated in order, with each step conditioned on all previously generated content and relevant literature, ensuring coherence and consistency. 3) References Integrity. To address the low reference count and citation formatting errors in LLM-authored papers, we implement a citation module. The research agent integrates both provided (e.g., retrieved literature) and autonomously generated references (e.g., models, datasets, foundational works) in the writing process. For each generated reference, search keywords are extracted from its context and used to locate the original work via Google and arXiv. References that cannot be reliably verified are removed, ensuring the integrity of the bibliography. 4) Incremental Compilation. After each section, the manuscript is compiled to detect and fix errors early. The final manuscript is compiled into a PDF. The prompt used for paper generation is provided in Appendix A.1.1, and an example of the generated paper can be found in Appendix A.2.

We build our review agent on top of AgentReview (Jin et al., 2024). Given a paper in PDF format, the system processes it through a five-stage pipeline simulating the peer-review workflow: (1) Reviewer Assessment I, (2) Author–Reviewer Discussion, (3) Reviewer Assessment II, (4) Meta- Review Compilation, and (5) Final Decision.

Discussion. In our analysis of Section 6, we observed that research deemed valuable by expert human reviewers was systematically undervalued by LLM-based reviewers, providing another clue to a potential source of bias. In particular, lower-scoring human-authored papers disproportionately addressed critical topics such as biases, risks, adversarial attacks, and limitations. Using keyword-based detection, we identified such papers and computed the sentiment polarity of their abstracts. Within the humanauthored set, the frequency of negative keywords was positively correlated with sentiment polarity, yet negatively correlated with review scores. Conversely, in LLM-authored papers, sentiment polarity showed no clear trend but remained consistently positive regardless of the number of negative keywords, and review scores increased with higher counts of such keywords. These findings suggested that a negative framing of critical topics could exacerbate bias in LLM reviews, whereas a positive framing of similar topics tended to yield disproportionately higher scores.

Takeaway (1) As the use of LLM-based polishing is permissible in most conferences, the tendency of LLM reviewers to assign inflated scores to submissions exhibiting LLM-generated stylistic features raises substantial fairness concerns for their practical deployment. (Section 6 & 7.1) (2) Research addressing bias, fairness, limitations, and other negative topics, tends to be systematically undervalued by LLM reviewers. (Section 7.2). (3) Submitting LLM-authored papers to academic venues wastes scholarly resources and undermines the integrity of peer review. Linguistic features can serve as preliminary indicators of such authorship. (Section 7.1) (4) Revisions guided by LLM review and revise yield quality gains in both LLM-based and human evaluations, illustrating the potential of the LLMs-as-reviewers paradigm to support early-stage researchers and enhance low-quality papers (Section 6).

Conclusion. Our multi-round simulation of LLM-driven research and review processes reveals systematic biases when LLMs are served as reviewers. LLM reviewers tend to overrate LLM-authored papers, disproportionately reward revisions, and undervalue critical human-authored work, leading to marked misalignment with human judgments. These biases are rooted in linguistic feature preferences and framing effects in critical discussion. Our findings highlight that, despite promising capabilities, LLM reviewers cannot yet be fully trusted as impartial evaluators in the scholarly ecosystem, especially in scenarios where they assess LLM-generated research. Addressing these biases is essential for the integrity and fairness of future AI-integrated scientific workflows.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do restrictions on reviewer LLM use actually shape peer review behavior? What prevents LLMs from applying their reasoning knowledge to improve outputs? How can we detect and account for LLM involvement in academic writing? Do language models reason through disagreement or only accommodate it? How can we reduce inherent biases in LLM-based evaluation judges? What limits language model accuracy in evaluating ideas? Why do LLM research ideation systems generate novelty but lack diversity?