How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Paper · arXiv 2608.08975 · Published August 10, 2026
LLM Evaluations and Benchmarks

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer’s original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains.

Introduction. Large language models (LLMs) are increasingly involved on both sides of scientific evaluation: they can revise how authors present their work and serve as scalable evaluators that accelerate review and reduce reviewer burden (Wang et al., 2020; Liang et al., 2024; Thakkar et al., 2026; Chen et al., 2026; Kaneko, 2026; Baumann et al., 2026). As these uses meet in the same evaluation pipeline, a central concern is how rhetorical presentation can reward-hack AI reviewers by changing their judgments without corresponding improvements in the underlying science. Presentation is especially relevant because the same scientific work can be communicated through different rhetorical choices: supported claims may be framed more assertively or cautiously, reported evidence may receive greater or lesser emphasis, and technical content may be expressed with different levels of complexity.

Discussion / Conclusion. This paper shows that AI scientific review is systematically sensitive to rhetorical presentation even when reported scientific content is preserved. This sensitivity is not uniform. It depends on the rhetorical dimension, the rewriting process, the rewriter and reviewer models, and the review protocol. The resulting variation cannot be reduced to a single average effect or treated as a stable property of one model configuration. These findings suggest that AI-assisted review should be evaluated for rhetorical robustness across multiple models and conditions, rather than judged solely by aggregate agreement or average scoring behavior. This study has several limitations. First, the benchmark is restricted to ICLR 2026 submissions with recoverable full-paper source and public review metadata, so the results may not generalize to other venues or scientific fields. Second, full-paper rewriting and multi-model evaluation are computationally expensive.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do we evaluate AI systems when user perception misleads actual performance? Why can't humans reliably detect AI-generated text despite measurable linguistic signatures? How do evaluation mechanisms prevent error accumulation in autonomous research systems? Does AI text rewriting systematically distort writer intent and preference? How should human oversight be integrated with autonomous AI systems? Why does verification consistently lag behind AI generation? When should tasks involve human-AI partnership versus full automation? Does conversational format create illusions of genuine AI communication? Do language models learn genuine linguistic structure or just surface patterns? Why do language models struggle with implicit discourse relations?