INQUIRING LINE

Can software tell when AI helped write a science abstract just by reading the text, or do light edits blend right in?

Can text-based algorithms reliably detect LLM assistance in scientific abstracts?

This explores whether software that reads only the text of a scientific abstract can reliably tell when an LLM helped write it. The corpus has little direct evidence on detection algorithms themselves, but it has useful material on why text alone is weak evidence and on what institutions do instead.


This explores whether software that reads only the text of a scientific abstract can reliably tell when an LLM helped write it. The corpus does not benchmark detection algorithms directly, so it can't give you accuracy numbers. What it does show is that the text itself carries a weak signal. When ML-literate readers tried to sort LLM-written from human-written abstracts, they couldn't do it reliably and leaned toward assuming a human was involved every time Can readers tell LLM abstracts from human ones?. The detail worth noticing is that LLM-edited abstracts got the highest clarity ratings and were preferred 55% of the time even when readers knew how they were written. 'LLM assistance' is therefore not one thing to catch. Much of it is light editing that makes a human draft better, and it blends in.

The problem may get harder as models improve. A study of how LLMs damage documents during editing found that weaker models leave visible traces, mostly by deleting content, while frontier models make subtle changes that keep the surface looking intact Does model capability change how documents degrade?. That study was about errors, not authorship. Still, the pattern carries over: the better the model, the less its output looks like a model wrote it. A text-based detector is chasing a target that keeps getting smoother.

The most practical evidence comes from ICLR 2026. Its program chairs used LLM detectors but did not trust them to make decisions. Detector flags went to area chairs as one input among several, and multiple human review steps were there to absorb false positives How can conferences detect and handle LLM misuse in peer review?. The one thing they enforced firmly was something they could check: references that turned out to be fabricated led to desk rejection. They stopped asking 'does this text sound machine-written?' and asked 'does this claim check out?' A hallucinated citation can be verified. A stylistic hunch can't.

A side lesson: if you use an LLM as the detector or judge, it brings its own surface biases. LLM judges are reliably swayed by fake authority signals and polished formatting, with no access to the model needed Can LLM judges be fooled by fake credentials and formatting?. Human readers fall for a similar shortcut, trusting answers with more citations even when the citations are irrelevant Do users trust citations more when there are simply more of them?. Any system that judges text by its surface, whether human or machine, can be gamed by changing the surface.

The corpus supports a 'no, not reliably,' with a useful reframe. The workable question is not 'did an AI touch this?' It is 'is anything here false or made up?' If you specifically want how well dedicated detectors perform on abstracts (true and false positive rates, robustness to paraphrasing), this collection doesn't yet cover it.


Sources 5 notes

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.