Can you reliably catch one specific paper or peer review as AI-written, or can only the overall share be estimated?
Can researchers detect individual papers modified by LLMs reliably?
This explores whether anyone, whether human readers, automated detectors or conference organizers, can reliably flag a specific paper or review as LLM-written or LLM-edited, as opposed to estimating how common LLM use is across a whole field.
This explores whether a single paper or review can be reliably caught as LLM-modified. The corpus mostly says no, at least not one document at a time, and the most credible measurements get around the problem by not trying. The best-known estimate, that 6.5% to 16.9% of review text at ICLR, NeurIPS, CoRL and EMNLP was substantially LLM-modified, comes from a distributional method. It measures shifts in word usage across thousands of reviews and estimates what share of the whole collection was touched by an LLM. It never labels any individual review (How much peer review text shows signs of LLM modification?). The method was designed to answer how much LLM use there is because the per-document question was too unreliable to answer.
Humans don't fill the gap. Readers with machine-learning expertise could not reliably tell LLM-generated research abstracts from human ones and tended to assume a human was involved whatever the source. LLM-edited abstracts also got the highest clarity ratings (Can readers tell LLM abstracts from human ones?). So the editing that is hardest to spot is also the editing readers like best. Work outside peer review points the same way: frontier models tend to change documents through subtle corruption that leaves the surface intact, while weaker models visibly delete content (Does model capability change how documents degrade?). As models improve, their fingerprints get fainter.
Conferences have adjusted to this. ICLR 2026 treated detector flags as one signal sent to human area chairs, not as automatic verdicts, because false positives were a real risk. The violation it did enforce firmly was fabricated references, which can be checked against the literature without guessing who wrote what (How can conferences detect and handle LLM misuse in peer review?). The broader idea is to stop asking 'did an AI write this?' and ask 'is something here demonstrably false?' ICML 2026 ran a randomized experiment that either banned or limited LLM use in reviewing. Substantial fractions of reviewers broke whichever rule they were given, and the rule barely changed scores or decisions (Does banning LLM use in peer review change review outcomes?). A policy that depends on detection only works if the detection works.
One result is unexpected: models may be better at recognizing their own writing than people are. Fine-tuning LLMs to recognize their own summaries increased how much they preferred those summaries, roughly in step (Do LLMs favor their own text because they recognize it?). That suggests some detectable signal exists, though the evidence is preliminary. It also suggests a risk: that signal can show up as bias in AI judges. Population-level analysis has its own pitfalls. Across 125,000+ reviews, LLM-assisted reviewers seemed to favor LLM-assisted papers, but the effect disappeared once paper quality was controlled for (Do LLM reviewers actually favor LLM-written papers?). Even aggregate detection needs careful statistics before anyone concludes something.
A caveat on coverage: most of this material is about peer reviews and abstracts, not full papers, and the corpus has no head-to-head benchmark of per-document detectors. What it supports is a practical lesson. Field-wide estimates are reliable, individual verdicts are not, and the workable enforcement targets are checkable errors like invented citations, not authorship itself.
Sources 7 notes
Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Show all 7 sources
Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Stop Automating Peer Review Without Rigorous Evaluation
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- A Retrospective on the ICLR 2026 Review Process
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights