INQUIRING LINE

Nobody has measured how often journal editors catch obvious text problems, but the documented misses show where they slip through.

How often do journal editors catch obvious textual problems before publication?

This explores how well the editorial and peer-review gate catches visible text problems, such as strange phrasing, fabricated references, hidden instructions or AI-generated passages, before a paper is published. The corpus has no direct catch rate, but it has several case studies showing what slips through and why.


This explores how reliably journal editors and reviewers catch obvious text problems before publication. The honest answer first: none of these notes measures an overall catch rate for editors. What the corpus does have is a run of documented misses, and they share a pattern. Problems that look obvious afterward got through when review was rushed, when the problem couldn't be seen in the format reviewers read, or when it was polished enough to look normal.

The clearest case is a single Elsevier journal, Microprocessors and Microsystems. Its 2021 volumes contain clusters of 'tortured phrases', the odd synonym swaps that automated paraphrasing tools leave behind. Those clusters line up with abruptly shortened review timelines and suspicious submission patterns Did automated text tools produce suspicious phrases in one journal?. The phrases were plainly odd, and they were still printed. So the useful question may be less about how good editors are and more about how much time they're given. A related study of AI conference peer review found the same link to rushing: reviews that showed heavy LLM modification came more often from low-confidence, rushed and less-engaged reviewers How much peer review text shows signs of LLM modification?.

Some problems aren't obvious at all to the people doing the reading. Researchers at institutions in eight countries hid invisible text in their papers instructing AI reviewers to be flattering. Standard PDF readers don't display it, but AI systems processing the file do read it Are researchers hiding prompt injections in academic papers? Are hidden AI prompts in preprints a deceptive research practice?. Readers with machine-learning expertise also couldn't reliably tell LLM-written abstracts from human ones Can readers tell LLM abstracts from human ones?. Sakana AI's fully AI-generated paper scored above the acceptance threshold in a double-blind workshop review. The citation error in it was found later by its own authors, not by the reviewers Can AI-generated papers pass peer review undetected?. Upstream of review, writers edited AI-drafted paragraphs only 23% of the time, and their edits left the text about 96% the same Do writers actually edit AI-generated text before publishing?. Much of what reaches an editor has therefore already skipped one round of human checking.

As models improve, this problem may get worse rather than better. Weaker models damage documents in visible ways, by deleting content. Frontier models tend to corrupt content quietly while the surface still looks intact Does model capability change how documents degrade?. The errors that make it to an editor's desk are increasingly the ones built to look fine.

The corpus also shows what does work. ICLR 2026 did not treat imperfect AI-text detectors as automatic filters. It sent their flags to human area chairs and desk-rejected only papers with confirmed fabricated references, because a fake citation can be checked How can conferences detect and handle LLM misuse in peer review?. In the other direction, an agentic AI reviewer that checks proofs line by line found critical errors in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. That gives a rough answer to the original question: human reviewers miss even serious errors often enough that a machine can find them in accepted papers. The workable approach seems to be the one ICLR used, where each check is matched to the kind of problem it can actually verify.


Sources 10 notes

Did automated text tools produce suspicious phrases in one journal?

Analysis of 1,078 articles in Microprocessors and Microsystems identified nonstandard phrases clustered in 2021 volumes with abruptly shortened review timelines, suspicious submission patterns, and partial detector flags suggesting possible automated text generation.

How much peer review text shows signs of LLM modification?

Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.

Are researchers hiding prompt injections in academic papers?

Researchers at institutions across eight countries embedded invisible instructions in paper PDFs and HTML telling AI models to produce flattering summaries. The text remains undetectable in standard PDF readers but executes when AI systems process the full document.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Show all 10 sources
Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.