INQUIRING LINE

Could authors slip hidden instructions for AI reviewers into a review copy, then scrub them out before publication?

Could hidden prompts be inserted during review and removed before publication?

This explores whether authors could slip hidden instructions aimed at AI reviewers into the version of a paper sent for review, then delete them from the published version, and what the collection says about how hard that would be to catch.


This explores whether authors could slip hidden instructions aimed at AI reviewers into the review copy of a paper and then delete them before publication. The collection doesn't study that exact switch, but it shows every piece needed for it to work. Hidden prompts are already in real papers. A Nikkei Asia investigation found invisible text in papers from at least 14 institutions in eight countries. The text was white-on-white or tiny, so it didn't show up in a normal PDF reader, but an AI that read the full document would see it and be told to write a flattering summary Are researchers hiding prompt injections in academic papers?. A separate audit found 18 arXiv manuscripts with hidden instructions telling AI reviewers to give positive assessments Are hidden AI prompts in preprints a deceptive research practice?.

The detail that answers your question is how these were caught: they were still in the public versions. That suggests the authors who were caught didn't bother removing them. A more careful author who put the prompt only in the submitted copy would leave nothing for a later scan of public papers to find. The arXiv audit also makes an ethics point that applies directly. The authors call it a questionable research practice because the text is both concealed and self-serving, whatever the author claims they meant by it Are hidden AI prompts in preprints a deceptive research practice?. Removing the prompt before publication would be two layers of concealment, not a defense.

The bigger question is whether hidden prompts are even needed. One study found that AI reviewers fail the two basic conditions for automating peer review. They agree with each other more than human reviewers do, a 'hivemind' effect. And they are easy to game: simply rewriting a paper's text with an AI raised AI review scores by about 0.45 points without improving the science at all Can AI systems safely replace human peer reviewers?. Hidden prompts are just the crudest version of a broader weakness. If AI reviewers respond to how a paper is worded, someone can steer them through what's visible as well as what's hidden. And because the reviewers tend to agree, one trick that works on one may work on all of them.

From other parts of the collection, two kinds of defense appear, though none was built for this problem. The first is to make AI review more rigorous. An agentic reviewer that checks proofs and experiments line by line caught serious flaws that had passed human review at major conferences Can inference scaling help reviewers catch errors humans miss?. At ICLR 2025, AI was used to give feedback to human reviewers, and 27% of them revised their reviews; the human stayed the final judge Can LLM feedback help peer reviewers improve their own reviews?. The second is to make documents traceable. Work on AI 'skills' argues that anything shaping an AI's behavior should be kept as versioned files that can be inspected and rolled back, not as hidden prompt state Can person-grounded skills remain auditable without hidden prompt state?. Applied to peer review, the fix for insert-then-remove is to archive the exact file reviewers saw and compare it with the published version. Writers in shared editors also said they want to see when and how AI was used Do writers want to see each other's AI prompts in shared editors?, which points toward disclosure as a norm.

What you might not have expected: the risk isn't only that a trick slips through. Sakana AI's fully AI-generated paper cleared a workshop's double-blind review with a 6.33 score, and its authors later found a citation error in it Can AI-generated papers pass peer review undetected?. The bigger danger is a review process where AI writes the papers, AI reviews them, and the review copy differs from the published one. In that setup, nobody is checking the version that actually matters.


Sources 8 notes

Are researchers hiding prompt injections in academic papers?

Researchers at institutions across eight countries embedded invisible instructions in paper PDFs and HTML telling AI models to produce flattering summaries. The text remains undetectable in standard PDF readers but executes when AI systems process the full document.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Show all 8 sources
Can person-grounded skills remain auditable without hidden prompt state?

COLLEAGUE.SKILL treats distilled expertise as versioned files subject to inspection, correction, and rollback—not hidden prompt state. Separating capability tracks from behavior tracks enables independent audit of what someone knows versus how they act.

Do writers want to see each other's AI prompts in shared editors?

Sixteen paired writers showed strong preference for higher levels of prompt visibility in shared editors, valuing awareness of when, how, and where AI was used. Benefits included understanding collaborators' thinking and verifying AI-generated text, though some found full sharing intrusive and self-conscious.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.