INQUIRING LINE

Polished AI text can fool human reviewers more easily than the detectors built to catch it, because they look for different things.

Can polished AI text fool both reviewers and detection methods?

This explores whether AI-written text that reads smoothly and looks professional can get past two kinds of gatekeepers: the humans and AI systems who judge its quality (like peer reviewers), and the tools built to flag it as machine-made.


This explores whether polished AI text can get past both the people and AI systems who judge its quality and the tools meant to flag it as machine-made. In this collection the short answer is: reviewers are easier to fool than detectors. That isn't because detectors are strong. It's because what fools a reviewer and what a detector looks for are different things.

People do badly at this. A review of 30 studies found that human judgment of whether content is AI-made sits around chance, and it hasn't kept up as AI output got more realistic Can people reliably spot content made by AI?. Peer review shows what happens next. One of Sakana AI's fully AI-generated papers scored above the acceptance line in double-blind review at an ICLR workshop. Its own authors later found a citation error and judged none of their three submissions good enough for the main conference Can AI-generated papers pass peer review undetected?. So the paper got through on appearance, not on soundness. One reason this works: polished output borrows an old shortcut, 'professional-looking means expert-made.' Less experienced readers, who can't check the substance underneath, are the most exposed Does polished AI output trick audiences into trusting it?. In an 81-person study, readers given no source signals couldn't tell fluent fabrications from true statements at all. They regained that ability only when shown which claims had been verified Can readers tell truth from fabrication without evidence signals?.

AI reviewers are, if anything, easier to fool. LLM judges reliably give higher scores to answers that include fake references or rich formatting, whatever the content says. These attacks need no access to the model and no optimization Can LLM judges be fooled by fake credentials and formatting?. For peer review in particular, a simple rewrite of a paper's text raised AI review scores by 0.45 points without changing the science. AI reviewers also agree with each other more than human reviewers do, so a trick that works on one probably works on all of them Can AI systems safely replace human peer reviewers?. Some authors have already acted on this: 18 arXiv manuscripts contained hidden instructions telling AI reviewers to be positive Are hidden AI prompts in preprints a deceptive research practice?.

Detection is where the picture gets more hopeful. Heavy rewriting does make text sound less like any one author. But the claim that it also beats AI-text detectors comes from a paper that never ran a detector, so that 'double erasure' is still unproven Do rewrites that hide authorship also fool AI detectors?. In fiction, a classifier that ignored style entirely still separated AI stories from human ones with 93% accuracy. It looked at narrative choices instead, such as how much characters act on their own and how events are ordered in time. Polishing the surface doesn't change those choices; only a real rewrite does Can AI stories be detected without analyzing writing style?. On the review side, an agentic reviewer that spends extra compute checking proofs and experiments line by line found serious flaws in papers that had passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?.

The takeaway you might not expect: polish defeats any judge that relies on surface impressions, whether human or AI. It does little against checks that look at structure or verify the substance. The bigger risk may come before any reviewer sees the text. Writers edit AI-drafted paragraphs only 23% of the time, and those edits leave the text about 96% unchanged, so most AI prose reaches readers with almost no human filtering Do writers actually edit AI-generated text before publishing?.


Sources 11 notes

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Does polished AI output trick audiences into trusting it?

Generative AI produces visually sophisticated outputs without underlying judgment, leveraging the historical heuristic that professional-looking work signals expert thinking. This substitution is especially risky for less experienced workers who lack domain knowledge to evaluate substance beyond form.

Can readers tell truth from fabrication without evidence signals?

In an 81-person study, participants given no provenance cues showed no significant truth discernment (p = .43), falling for fluent hallucinations as readily as ground truth. An idealized Provenance Density interface showing verified claims restored a +4.15 point gap (p < .001).

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Show all 11 sources
Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Do rewrites that hide authorship also fool AI detectors?

The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.

Can AI stories be detected without analyzing writing style?

StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.