LLM or Human? Perceptions of Trust and Information Quality in Research Summaries

Paper · arXiv 2601.15556 · Published January 22, 2026
Expertise in the Age of AI Content

Large Language Models (LLMs) are increasingly used to generate and edit scientific abstracts, yet their integration into academic writing raises questions about trust, quality, and disclosure. Despite growing adoption, little is known about how readers perceive LLM-generated summaries and how these perceptions influence evaluations of scientific work. This paper presents a mixed-methods survey experiment investigating whether readers with ML expertise can distinguish between human- and LLM-generated abstracts, how actual and perceived LLM involvement affects judgments of quality and trustworthiness, and what orientations readers adopt toward AI-assisted writing. Our findings show that participants struggle to reliably identify LLM-generated content, yet their beliefs about LLM involvement significantly shape their evaluations. Notably, abstracts edited by LLMs are rated more favorably than those written solely by humans or LLMs. We also identify three distinct reader orientations toward LLM-assisted writing, offering insights into evolving norms and informing policy around disclosure and acceptable use in scientific communication.

Introduction. In early 2025, a PhD student at the University of Minnesota was expelled after being accused of using a Large Language Model (LLM) to generate answers for an exam [15]. The university’s determination rested on faculty reviewers’ judgments that parts of the student’s exam answers showed stylistic and structural similarities to answers produced ∗This work was completed while the author was at Amazon AWS AI.

Authors’ Contact Information: Nil-Jana Akpinar, niljana.akpinar@gmail.com, Microsoft, USA; Sandeep Avula, sandeavu@amazon.com, Amazon AWS AI, USA; CJ Lee, cjlee@amazon.com, Amazon AWS AI, USA; Brandon Dang, dangbran@amazon.com, Amazon AWS AI, USA; Kaza Razat, razat@amazon.com, Amazon AWS AI, USA; Vanessa Murdock, vmurdock@acm.org, Amazon AWS AI, USA.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org.

© 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM 1 arXiv:2601.15556v1 [cs.CY] 22 Jan 2026 2 Akpinar et al. by ChatGPT. The student, however, maintained that he had only used AI tools for grammar checking and denied generating any exam content with ChatGPT. The case drew national headlines not only because of the severity of the punishment, which is striking given that detecting LLM-generated text is far from certain, but also because it reveals a broader tensions surrounding AI use in academia and scientific communities. As LLMs become increasingly capable at complex tasks like summarization [2, 60], paraphrasing [53], and translation [40], more scholars are adopting them into their research processes [30, 43]. A 2024 survey of 816 academic authors across disciplines and found that 81% had already used LLM tools in some aspect of their workflow [31]. Analyses of recent OpenReview data suggest that between 6.5% and 16.9% of submitted peer reviews may already use LLMs [29, 63]. Yet, there is no consensus on what constitutes appropriate use and how, if at all, LLM use needs to be disclosed. Policies are mixed, with many publication venues prohibiting blind text generation (e.g., AAAI, AIES), some explicitly permitting LLM use for editing or polishing (e.g., CHI, CSCW, Neurips), and others not addressing the issue at all. The lack of consensus in policies is particularly salient when it comes to research summaries, which are both widely read and increasingly subject to LLM involvement.

Research summaries, and especially abstracts, are central to how scholarship is evaluated, shared, and cited. They function as the primary entry point for readers deciding whether to engage with a paper and are expected to be trustworthy, concise, accurate, and comprehensive. As LLMs become increasingly capable of generating or editing abstracts, concerns about how such texts are perceived and trusted become increasingly important.

At the same time, research summaries written with LLM assistance hold considerable promise. If they can be made trustworthy, they could broaden access to science by producing lay summaries for students, policymakers, and the public.

They may also support interdisciplinary understanding, reduce linguistic barriers through multilingual summarization, and reduce the workload for authors drafting abstracts. Realizing this potential requires a careful understanding of how readers perceive LLM-generated summaries and how those perceptions shape judgments of quality and trust.

Researchers have documented factual errors, stylistic vagueness, and overgeneralization in LLM-generated summaries [14, 41, 50]. Yet we still lack systematic evidence on how readers themselves perceive and trust such summaries in academic contexts. Hadan et al. [17] made some leeway on this question by studying reviewers’ ability to identify and evaluate AI-generated writing in the peer-review process, but this remains a single recent study focused specifically on reviewers, highlighting the need for additional investigation into general reader perceptions. More research is required to understand whether readers can reliably distinguish LLM- from human-written abstracts, and how actual versus perceived involvement of LLMs shapes judgments of quality and trust. Moreover, little is known about the orientations that readers adopt toward LLM-assisted writing and how these stances manifest in evaluation practices.

Related work. 2.1 Trust, Information Quality, and Measurement Trust is often described as “the attitude that an agent will help achieve an individual’s goals in a situation characterized by uncertainty and vulnerability” [27]. In the context of AI, trust captures users’ attitudes toward relying on model outputs under uncertainty, distinct from reliance (behavior) and trustworthiness (system quality).

Studies of AI-assisted writing find that language models improve surface quality such as grammar and fluency but may reduce depth or originality [17, 19]. These mixed effects make trust a key outcome for understanding how readers evaluate AI-authored or AI-edited text.

Trust has been measured through behavioral and self-report metrics, including reliance calibration and weight-ofadvice approaches [35, 55–57]. Across studies, performance and perceived quality are among the strongest predictors of trust [58]. Transparency and explanations can strengthen trust but show mixed effects on calibrated reliance [7].

Individual differences also matter: users with more prior exposure to AI often exhibit higher trust, whereas those with deeper technical expertise may be more cautious [4]. To answer RQ2 and RQ3, we evaluate abstracts along established dimensions such as clarity and comprehensiveness, aligning our measures with prior work on trust and information quality [33, 38, 44].

2.2 Trust in LLM-Generated Text and Summaries Recent work highlights concerns for trust in LLM-generated summaries. Huang et al. [22] establishes a benchmark for trustworthiness, emphasizing that it is multi-dimensional, involving not only accuracy but also robustness and fairness. Several empirical studies have evaluated LLMs directly in summarization tasks [59, 62]. In medical evidence summarization, Tang et al. [50] find that models can produce factually inconsistent summaries, sometimes overstating results and other times introducing ambiguity that misrepresented the evidence. Gao et al. [14] test ChatGPT on generating medical abstracts from only titles and journal-style cues. They show that while AI detectors could identify Manuscript submitted to ACM 4 Akpinar et al. most outputs, human reviewers misclassified a substantial fraction and described them as vaguer and more formulaic.

In academic writing, Peters and Chin-Yee [41] compare LLM summaries of full abstracts and articles to original texts and expert digests. They find the LLMs’ tendency to exaggerate claims and overgeneralize. Although this study did not measure trust directly, such amplification poses a risk that readers could accept overstated results as faithful representations. Beyond factual quality, trust is also shaped by perceptions: Hadan et al.

Method. 3.1 Study Design We designed an online survey-based experiment in which participants complete a series of three tasks evaluating three abstracts per task. Within a task, abstracts correspond to the same research paper but differ in authorship with the following options: the original human-written version (human-written abstract), a version generated from the full research paper by an LLM (LLM-generated abstract), and a version edited by an LLM based on the original human-written abstract (LLM-edited abstract). Each task uses one of the abstract types as the focal target of evaluation while the other two versions are used for comparison. The order of abstract types is counterbalanced using a Latin square design [16] to mitigate ordering effects and provide within-subject variation.

To assess how awareness of LLM involvement influences peoples’ judgments, we add a between-subjects component to the experiment which ultimately contributes to a mixed study design. Participants are randomly split into two conditions: An information condition, where the authorship of each abstract is disclosed, and a guess condition, where participants are asked to gauge the degree of LLM involvement without additional information.

Across all tasks and conditions, participants are asked to rate abstracts across various Likert scales related to information quality and trust. Participants are prompted to explain their choices in several free-text fields and, when presented with all three abstract types at once, are asked to identify their preferred abstract for each paper. Along with the main task-specific information, we also ask participants to provide feedback via post-task and post-survey questionnaires.

For each paper, we consider three abstract types. First, the original abstract created by the authors (Human-written).

Second, an LLM generated abstract that is written entirely by the model (LLM-generated). Third, an LLM rewritten abstract in which the model edits the original human-written abstract (LLM-rewritten. These three types represent realistic use cases from fully automated summarization to more collaborative forms of writing.

For LLM-rewritten abstracts, the model receives the original human written abstract together with the prompt:

Abstract guidelines: [ACM INSTRUCTIONS] Rewrite this Abstract for a scientific paper with title ’[PAPER TITLE]’ based on the given abstract guidelines. Assume you are the author of the paper and start with ’Abstract:’.

Abstract: [ORIGINAL ABSTRACT] 3.4 Measures and Survey Procedure 3.4.1 Procedure. Participants complete an online survey consisting of three main tasks. In each task, they are exposed to a combination of paper title and abstract, and answer a series of questions designed to assess their perceptions of LLM involvement, quality, and trustworthiness. Depending on the condition assignment, participants either evaluate Manuscript submitted to ACM 8 Akpinar et al. the abstract without knowing the extend of LLM involvement (guess condition) or they are told explicitly which of the three types of abstracts is shown to them (information condition). We then reveal all three abstract types for the same paper and solicit participants’ preferences and the reasoning behind them.

We conducted two rounds of pilot testing with internal participants. The pilots helped us to understand (1) whether the instructions are easily understandable, (2) whether the flow of the survey feels natural, (3) how long the survey takes and how many tasks each participant can complete, and (4) whether individual wordings of questions are clear.

The final survey was designed to take approximately 60 minutes, and participants were asked to refrain from using LLMs.

Discussion. 5.1 Findings in Context Our experiment examines how readers engage with research abstracts under varying degrees of LLM involvement. We first find that participants do not reliably distinguish between LLM-generated, human-written, or LLM-edited abstracts, tending instead to assume some degree of human involvement (RQ1). At the same time, they show a baseline suspicion that LLMs were involved across all abstracts. Participants’ judgments draw on heuristics such as completeness, clarity, credibility, engagement, and writing conventions, but these cues prove systematically unreliable. For instance, P31 praised LLM-generated content as written by "a well educated researcher who has a clear understanding of the problem," while P12 criticized a human-written abstract for "unnecessarily complicated phrases (most likely generated by an LLM)." These findings align with prior work surfacing readers’ inability to distinguish LLM and human-written texts in contexts like news articles, story writing, and peer review [9, 17, 24, 42]. Similar to our observations, Jakesch et al.

[24] describe that their participants used intuitive but often flawed heuristics to detect whether text is LLM-generated.

Porter and Machery [42] report on a human preference bias similar to ours where readers tend to judge LLM-generated works as human more often than the reverse.

For the next research question, we find that participants’ evaluations of abstracts reflect both the type of authorship and whether LLM involvement was disclosed. If authorship was disclosed, participants’ gave higher quality ratings to human-written and LLM-edited abstracts as compared to LLM-generated ones. Quantitatively, LLM-edited abstracts received the highest clarity ratings (β= 1.383, p< .001) and were selected by 55% of participants when authorship was disclosed, compared to 27-28% for human-written and LLM-generated alternatives. Participants valued that LLM editing achieved clarity without sacrificing substance. P15 noted that "LLM introduced linguistic clarity and cohesiveness," while P26 praised abstracts that "strike a balance between clarity and technical depth." By contrast, LLM-generated abstracts were criticized for "information overload without focus" and including "too many specific terms from the paper" (P4), while human-written abstracts were valued for coherent structure but sometimes lacked quantitative support. Readers’ preferences for text that is human-written or written in collaboration of human and LLM, rather than entirely generated by an LLM, has been reported by past studies [11, 20, 23]. Across all abstract types, disclosure consistently elevated ratings of trust and quality. Participants expressed neutrality when guessing but reported positive trust once authorship was revealed; even for fully LLM-generated texts. Abstract preference selection reinforced this pattern, with LLM-edited versions strongly preferred, but only when their LLM authorship level is revealed. Together our findings suggest that while abstract type shapes specific strengths and weaknesses, disclosure of LLM involvement may have a more substantial impact on how readers judge trustworthiness and information quality. This finding contrasts the findings of prior work demonstrating decreases in credibility and quality judgments after AI use has been disclosed [3, 8, 32]. We discuss this phenomenon further in Section 5.3.

Moving to the perceived level of LLM involvement, our results show that participants’ beliefs about authorship shaped their evaluations even when inaccurate (RQ3), which is in line with prior oberveations from the literature [3, 13, 24].

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How reliably can humans and AI detectors identify machine-generated text? Does disclosing AI authorship change how audiences evaluate the writing? How can we detect and account for LLM involvement in academic writing? Can readers reliably distinguish AI-written text from human writing? Can AI systems perform peer review as effectively as humans? How do hallucinated citations emerge in AI scholarly output? Do restrictions on reviewer LLM use actually shape peer review behavior?