SYNTHESIS NOTE
Topics›Evaluations›this note

How much does rhetorical style shift AI review scores?

When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.

Synthesis note · 2026-09-25 · sourced from Evaluations

The paper's central finding is that "AI scientific review is systematically sensitive to rhetorical presentation even when reported scientific content is preserved." The authors frame this as a possible form of reward hacking: presentation changes an AI reviewer's judgment "without corresponding improvements in the underlying science." The sensitivity is not a flat penalty or bonus for polish. It is "structured rather than uniform." Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, scope framing is a weaker second tier, and the remaining dimensions have smaller or less stable effects.

The design supports a causal reading. From 120 anonymized ICLR 2026 submissions the authors built a controlled corpus of 4,200 full-paper manuscripts. Two LLM rewriters transformed six rhetorical dimensions in opposing directions, and five LLM reviewers scored the results under standard and strict protocols. Because the same paper appears in both a more favorable and a less favorable rhetorical version, the contrast isolates presentation from content. The hierarchy of dimensions persists across human-assessed quality levels. Score movement, though, depends strongly on the reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. The paper also tests joint, recursive and reviewer-guided rewriting and finds that "more elaborate workflows do not reliably yield larger gains."

This sits beside a group of notes about fluency standing in for substance. Do fluent arguments win debates through sound logic or rhetorical polish? shows LLMs judged well on persuasiveness and poorly on formal argument strength. Does polished AI output trick audiences into trusting it? makes the same point about human audiences reading polished artifacts. This paper adds a scientific-review setting in which the evaluator is itself an LLM. It also adds a finer resolution: which rhetorical levers move a score (how evidence is framed, how novelty is stated), not just that style matters. The closing recommendation is that AI-assisted review be checked for rhetorical robustness "across multiple models and conditions," not judged by aggregate agreement or average scoring behavior. That is a warning aimed at benchmarks like the alignment figures in Can structured pipelines make LLM novelty assessment reliable?, where agreement with humans says nothing about stability under rewording.

The excerpt leaves a lot unstated. It gives no effect sizes, does not name three of the six dimensions, and does not say which reviewer models or rewriters were used or how the strict protocol differs from the standard one. It does not explain why scores move toward the middle. The authors also flag that the benchmark is limited to ICLR 2026 submissions with recoverable source and public review metadata, so generalization to other venues or fields is open, and full-paper rewriting and multi-model evaluation are expensive. Nothing here tests whether an inference-scaled reviewer of the kind in Can inference scaling help reviewers catch errors humans miss? is exposed in the same way. The defensible reading is narrow: for the reviewers and protocols tested, rhetoric alone moves scores, and any claim about AI review quality should say under which rhetorical conditions it was measured.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What safeguards enable trustworthy AI-assisted scientific peer review at scale?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 127 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

AI review is systematically sensitive to rhetorical presentation with scientific content preserved — evidence framing and novelty stance move scores most