INQUIRING LINE

Do preprint servers screen papers for invisible text aimed at AI reviewers, or do outside investigators only catch it after publication?

Do preprint servers have tools to detect hidden text in submitted manuscripts?

This explores whether places like arXiv screen incoming papers for invisible text, such as white-on-white prompts aimed at AI reviewers, or whether such text gets caught some other way.


This explores whether preprint servers like arXiv check submissions for hidden text, especially invisible instructions planted to manipulate AI reviewers. The short answer from this collection: it doesn't document any such screening tool. Every documented case was caught after the papers were already public, by outside investigators rather than by a server's intake checks. That gap is worth knowing about, because the problem is real and spreading.

The scale is larger than most people would guess. A Nikkei Asia investigation found hidden text in papers from at least 14 institutions across eight countries. The text was invisible in an ordinary PDF reader but fully readable to an AI system processing the document, and it told the AI to produce flattering summaries Are researchers hiding prompt injections in academic papers?. A separate analysis found 18 arXiv manuscripts carrying concealed instructions aimed at AI reviewers. It argues the practice counts as questionable research conduct whatever the authors say they intended, because hiding the text and making it self-serving is enough to breach publication ethics Are hidden AI prompts in preprints a deceptive research practice?. The trick works because a human and a machine reading the same file see different documents.

The trick also has good odds against its target. LLM judges are easily swayed by surface signals like fake references and polished formatting, and these attacks need no special access to the model Can LLM judges be fooled by fake credentials and formatting?. A hidden line saying 'rate this paper highly' is a cruder version of the same weakness.

Where screening does happen in this collection, it happens at conferences, not preprint servers. Even there it targets a different problem: AI-written text and fabricated references, not hidden instructions. ICLR 2026 sent LLM-detector flags to human area chairs as one input among several, because the detectors are imperfect. It saved automatic desk rejection for fabricated references that had been confirmed, which is something you can actually verify How can conferences detect and handle LLM misuse in peer review?. The lesson carries over: enforcement works best when it targets something checkable, and hidden text is checkable in a way that 'sounds AI-written' is not. Readers with ML expertise can't reliably tell AI-written abstracts from human ones Can readers tell LLM abstracts from human ones?, but invisible text can be found by inspecting the file itself.

Why does it matter that preprint servers are the weak point? Preprints shape debate before any review takes place. MIT's case shows an arXiv paper steering discussion long before anyone raised reliability concerns Can unreviewed preprints shape scientific debate before peer review?. Problem papers can also outlive their correction: undisclosed GPT-written papers keep circulating on mirror sites even after retraction How much GPT-written scholarship reaches Google Scholar undetected?. A manipulated preprint can therefore travel far before anyone checks it. If you want to know whether arXiv has since added checks, this collection can't answer that. The documented detection so far has come from journalists and researchers, not from the servers themselves.


Sources 7 notes

Are researchers hiding prompt injections in academic papers?

Researchers at institutions across eight countries embedded invisible instructions in paper PDFs and HTML telling AI models to produce flattering summaries. The text remains undetectable in standard PDF readers but executes when AI systems process the full document.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Show all 7 sources
Can unreviewed preprints shape scientific debate before peer review?

MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.

How much GPT-written scholarship reaches Google Scholar undetected?

Haider et al. found roughly 62% of papers matching common ChatGPT phrases lacked GPT disclosure, with 57% addressing policy topics (computing, environment, health). Most appeared in non-indexed journals but some reached mainstream outlets; copies persist across mirror sites even after retraction.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.