Can readers tell LLM abstracts from human ones?
Do readers with ML expertise reliably distinguish human-written, LLM-generated, and LLM-edited research abstracts? Understanding this matters for evaluating whether readers can serve as effective gatekeepers against LLM content.
Akpinar et al. report that readers with machine learning expertise do not reliably separate human-written, LLM-generated and LLM-edited research abstracts. The excerpt says participants "struggle to reliably identify LLM-generated content," and the discussion finds them "tending instead to assume some degree of human involvement" in all three types, with "a baseline suspicion that LLMs were involved across all abstracts." The types were not treated alike, though. With authorship disclosed, LLM-edited abstracts "received the highest clarity ratings (β= 1.383, p< .001) and were selected by 55% of participants when authorship was disclosed, compared to 27-28% for human-written and LLM-generated alternatives."
The authors trace these judgments to heuristics they list as "completeness, clarity, credibility, engagement, and writing conventions," and they conclude that these cues "prove systematically unreliable." Readers valued LLM editing because it "achieved clarity without sacrificing substance"; one participant noted that "LLM introduced linguistic clarity and cohesiveness." LLM-generated abstracts drew criticism for "information overload without focus." The advantage was conditional: LLM-edited versions were "strongly preferred, but only when their LLM authorship level is revealed."
This is a reader-side case of a pattern in How much does rhetorical style shift AI review scores?, where rewriting presentation without changing reported content moved LLM reviewers' scores. The lever here is clarity editing and the readers are human, but both show presentation moving judgments of the same underlying science. The clarity preference also fits the account in Does polished AI output trick audiences into trusting it?, where polish stands in for expertise. The excerpt is consistent with that account but does not test whether readers checked substance. The disclosure half of the finding, which the same survey reports separately, is in Do reader judgments reflect actual authorship or just their beliefs?.
The excerpt does not give the number of participants, how they were recruited, or how often readers identified the true authorship type, so "do not reliably" is the strongest statement it supports; no accuracy rate appears in it. The 55% and 27-28% figures are selection shares under the disclosed condition of one survey, and the excerpt does not say what the β is measured against. Nothing here shows that LLM-edited abstracts are more accurate or more trustworthy than human ones; it shows that readers given these texts rate the edited version as clearer. At this strength, the implication is that a reader's sense of an abstract's quality is a weak check on whether an LLM was involved, so rules that expect readers to catch LLM text in summaries rest on thin ground.
Inquiring lines that read this note 49
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How reliably can humans and AI detectors identify machine-generated text?- Can readers distinguish machine-generated text from human-written comments?
- Do LLM detectors reliably identify generated text in real research settings?
- Do detector systems miss certain types of LLM-generated writing?
- Do readers who prefer LLM-edited abstracts check the substance or just clarity?
- Why does disclosure of LLM authorship change reader trust and preference?
- What heuristics do readers use to detect or fail to detect LLM writing?
- How do newer LLM generations differ from human writing patterns in detectable ways?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- Can readers reliably distinguish LLM-generated research writing from human writing?
- Does LLM use reduce writing costs differently across linguistic backgrounds?
- Can text-based algorithms reliably detect LLM assistance in scientific abstracts?
- How does LLM-modified writing narrow linguistic diversity in peer review?
- Why do computer science papers show more LLM modification than other fields?
- Can researchers detect individual papers modified by LLMs reliably?
- Did LLM use in scientific writing plateau after the initial surge?
- What concerns does widespread LLM use raise for scientific independence?
- How much do LLM reviewers shift scores based on rhetorical framing alone?
- What methods can reliably detect LLM-generated academic papers at scale?
- How does framing critical topics shape LLM review scores?
- Why do readers rate LLM-edited text more favorably?
- Do humans and LLMs agree on novelty assessment in research?
- How much does rhetorical framing shift LLM reviewer scores independent of content?
- Does exposure to LLM answers actually change how people think critically?
- How much does polished presentation substitute for actual expertise in reader judgment?
- Do texts judged as slop actually contain measurable stylistic patterns unique to LLMs?
- Can readers reliably distinguish AI-written abstracts from human-written ones?
- How do aggregate patterns in LLM text differ from what humans perceive?
- Does presentation style bias how evaluators judge scientific methods and results?
- What makes rhetorical polish misleading in evaluating research quality?
- Do preprint servers have tools to detect hidden text in submitted manuscripts?
- How often do journal editors catch obvious textual problems before publication?
- How do citation errors in AI-generated papers differ from human hallucinations?
- Do surface phrases reliably identify unedited machine-generated scholarship?
- Why are hallucinated references easier to detect and punish than LLM-assisted writing?
- How susceptible are LLM evaluators to fake references as exploitable biases?
- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Can rules against undisclosed LLM use change reviewer behavior without enforcement?
- Can watermark-based detection measure true prevalence of LLM use in peer review?
- Why does author withdrawal authority matter for preprint accountability?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do conference policies banning LLM use actually reduce AI involvement in reviews?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
- Can peer review policies actually prevent LLM use when compliance is hard to monitor?
- Do metareviewers and regular reviewers use LLMs differently in peer review?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
the same pattern in reviewers: presentation alone shifts judgments of unchanged scientific content.
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
polish standing in for expertise; this excerpt fits it without testing substance checks.
-
Do writers want to see each other's AI prompts in shared editors?
This study explores whether revealing AI prompting activity to collaborators in text editors affects how writers work together. Understanding prompt visibility matters because it shapes trust, learning, and awareness of AI's role in collaborative writing.
writer-side wish for visibility into AI use; this study measures the reader side.
-
Do reader judgments reflect actual authorship or just their beliefs?
When readers evaluate research abstracts, do their ratings track who actually wrote them, or are they shaped by what they believe about authorship—even when those beliefs are wrong?
sibling note from the same survey on belief and disclosure effects.
-
Does polished writing actually signal better quality work?
When evaluators judge applications and manuscripts, does rhetorical sophistication predict merit, or does it distract from verifiable evidence of competence and rigor?
Qualifies: rhetorical polish should not be read as merit, so the top ratings for LLM-edited abstracts may not indicate quality
-
Can people reliably spot content made by AI?
This systematic review of 30 studies asks whether human judgment can distinguish AI-generated text, images, and voice from human-created content, and whether detection accuracy has improved as AI becomes more realistic.
Evidence for: a 30-study review finds human accuracy at spotting generative AI text clusters near chance, across modalities
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Stop Automating Peer Review Without Rigorous Evaluation
- Scientific production in the era of Large Language Models
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- Mapping the Increasing Use of LLMs in Scientific Papers
Original note title
readers could not reliably tell LLM from human research abstracts, yet LLM-edited abstracts were rated most favorably