Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
This systematic review synthesized empirical evidence on human ability to distinguish generative artificial intelligence content from human-produced content across text, image, and voice modalities. A structured search of Scopus identified 22,541 records from 2025- 2026, of which 1200 were screened and 30 studies were included. Across these studies, human detection accuracy varied widely but generally clustered around chance performance. Overall, the literature shows that humans are generally unreliable detectors of gen AI content, raising broader questions about whether the ability to tell should matter for how we evaluate or trust content.
Introduction. Since the rise in popularity of Large Language Model (LLM) chatbots like ChatGPT or Google Gemini, also commonly referred to as artificial intelligence or AI, one of the most controversial and empirically studied topics in how people interact with these new technologies is the uncertainty people experience when judging whether a piece of content was produced by generative artificial intelligence (gen AI) [1–4]. This has only seemed to have accelerated, with at least 16 studies published in the first 3 months of 2026 alone [5– 19]. Diverse concerns motivate these studies, ranging from worries about the large-scale dissemination of potentially malicious AI-generated content such as social-engineering and disinformation mechanisms in online environments [7,20] to concerns about authenticity in personal application statements [12,14] and academic integrity in scientific manuscripts and essays [13,15]. Despite this proliferation of empirical work, there is yet to be a synthesis of how accurately humans can identify AI-generated content across modalities or contexts. To address this gap, this systematic review synthesizes current evidence on human ability to recognize different kinds of gen AI content.
Method. This review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines for study identification, screening, eligibility assessment, and inclusion. The overall study selection process is summarized in Figure 1.
A literature search was conducted in the electronic database Scopus, which was used as the primary source for indexed publications because it offers broad multidisciplinary coverage across fields such as computer science, psychology, linguistics, and human– computer interaction. Moreover, the database supports complex Boolean queries with wildcards, enabling better, context-specific retrieval of studies involving generative AI systems, human evaluators, and detection or discrimination tasks. The search was restricted to 2025–2026 to ensure the review reflects the most current evidence on human detection of generative AI content. The keywords used during the search included terms related to AI-generated media, human detection, and evaluation tasks. The final Boolean query was: (TITLE-ABS-KEY("AI generat" OR "machine generat" OR "synthetic text" OR "synthetic media" OR "large language model" OR "LLM" OR ChatGPT OR GPT OR "generative AI" OR "neural text generator")) AND (TITLE-ABS-KEY(detect* OR identify* OR distinguish* OR recogniz* OR "classification task" OR "human evaluation" OR differentiat* OR discriminat)) AND (TITLE-ABS-KEY(human OR particip OR user OR evaluator OR judge OR rater OR annotat* OR "lay person" OR "expert reviewer")) The review included published papers involving human participants who were asked to detect or distinguish AI-generated versus human-generated content across various modalities (text, image, voice). Eligible studies were required to report quantitative performance metrics from which a retrievable proportion of correct human classifications and the corresponding number of human evaluators could be obtained. Studies were not required to designate accuracy as the primary outcome, provided that sufficient data were available to compute a success proportion and sample size. Only studies published in English during 2025–2026 were considered. Studies were excluded if they were review articles, conceptual or theoretical papers, conference abstracts without results, or empirical studies that used AI-generated stimuli but did not measure human detection accuracy. Papers with incomplete information or insufficient methodological detail were also excluded.
The initial search yielded 22,541 records, ordered by relevance. Titles and abstracts were screened sequentially until 1,200 records were reviewed, at which point no eligible studies had been identified in the last 100 records, suggesting saturation. A total of 44 studies met the preliminary inclusion criteria and were retrieved for full-text evaluation. Following further evaluation, 14 studies were excluded for reasons such as lack of a human detection task (3), absence of extractable accuracy data (4), or general methodological ineligibility (7). A total of 30 studies were included in the final synthesis.
Data was extracted and summarized in Table 1. Extracted variables include publication details, modality of AI-generated content, population characteristics, sample size, and reported results involving human accuracy in detecting gen AI content. Where not explicitly provided, we used relevant quantities from the study data and results to obtain the average proportion of correct classifications and the corresponding number of human evaluators.
Discussion. Despite substantial methodological and contextual heterogeneity across the studies reviewed, a consistent pattern observed is that humans, on average and in general, do not distinguish AI-generated content from human-generated content reliably better than chance. This echoes earlier assessments and positions [35–37], suggesting that human accuracy has not meaningfully improved over time or at least has not improved at a pace that keeps up with the increasing realism of gen AI content [38].
Performance was observed in reviewed studies to vary by modality, with more success at identifying AI-generated voice content than images, and more success with images than with text. This is consistent with [39] which discussed how humans may be better at detecting synthetic voices than images or text due to subtle artifacts in timing, pitch, or prosody. For image detection, visual heuristics, such as inconsistent reflections, unnatural textures, or impossible lighting may help identify generative AI content [40], but accuracy rates from reviewed studies on images still mostly did not vary convincingly from chance, also consistent with previous studies [41].
In line with motivations for conducting these studies, an important subgroup of the reviewed literature focused on text content which has typically served as basis for evaluators’ judgments about the author’s expertise or competence. These include personal statements for applications [12,14,30] and academic manuscripts or essays [13,15,27,28]. Accuracy across these studies weighted by sample size is 58%. In addition, evaluators oftentimes perceived gen AI generated documents not only as human written, but of better quality [12,14,15]. It is important to point out that the reviewed studies’ prevailing discussion of their results is not to call on improving accuracy, but to challenge the prevailing use of how well these texts are written as a meaningful indicator of quality. Personal statements should only be valued insofar as they contain verifiable, objective information about a candidate’s experiences, decisions, and conduct. The candidate’s rhetorical sophistication should not be construed, explicitly or implicitly, as evidence of their suitability. Long before generative AI, applicants with access to professional editing or commercial statement-writing services were already advantaged by the same evaluative bias toward rhetorical refinement [12]. By lowering the cost of high-quality writing assistance, gen AI may in fact be partially leveling the playing field for applicants with fewer financial resources or lower linguistic ability, especially for those whom English is not a first language. Similar reasoning can be applied to scientific writing. The point of polished language in research manuscripts has always been merely to improve clarity. Scrutiny of the reasonableness of methods and the veracity of results is invariant to how the text was produced. As discussed in [42], gen AI may “serve as an equalizer” for researchers who struggle with academic English.
Naturally, this dynamic is different when generative AI is used with the intent to deceive, such as in fabricated customer reviews [23] or coordinated social-media manipulation campaigns [7]. In these contexts, the concern is not evaluative fairness but the risk that highly convincing and abundant AI-generated content can amplify misleading or harmful messages. The results of this review underscore the need for readers to adopt a more cautious, discerning stance toward online information, resisting the impulse to treat unverified claims as credible simply because they appear to be widely repeated.
Conclusion. Considerable human uncertainty in identifying gen AI content is real and is observable across modalities and, at least for text, contexts within the modality. This raises a broader question about what matters going forward: whether people can tell if something is gen AI content or whether the ability to tell should remain relevant to how we evaluate or trust content. Perhaps the better path is becoming more honest about the criteria we use to judge quality and more discerning about what we choose to believe from what we read, regardless of who or what actually produced the words.
Limitations. This review has limitations that should be considered when interpreting its findings. First, the search was conducted exclusively in Scopus, which, although broad in multidisciplinary coverage, may not index all relevant studies, particularly those in fast-moving computer science venues or preprint repositories. Second, while there were more than 20,000 hits from the initial search, only the first 1,200 ordered by relevance were screened. While the consistency of the findings across the reviewed studies suggests that the conclusions are stable, there may be studies missed that could provide further nuance. Third, substantial methodological heterogeneity across studies limited the ability to compute pooled effect sizes. Accuracy estimates therefore assume regularity across participants within each study, which may not reflect true variability in individual detection ability.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How reliably can humans and AI detectors identify machine-generated text?- How do lay readers differ from classifiers in detecting AI text?
- How much does the human-authorship halo affect AI evaluation across different task domains?
- What false-positive rates do AI detectors show on mixed human-AI drafts?
- Are detector errors on AI text systematic or random by design?
- Can AI detectors reliably distinguish human from machine-generated text?
- Can a classifier distinguish machine-written text from poor human writing?
- What signals do AI text detectors actually measure in their classification?
- Can detection systems identify AI text rewritten to match human author style?
- Would detectors trained on unaltered AI text catch heavily rewritten versions?
- Can AI-rewritten text still be detected as machine-modified?
- Can text detection methods distinguish between AI collaboration and delegation?
- Can user feedback flags rival AI detector accuracy for identifying AI slop?
- Why do consumers show lower valid-view rates for AI-generated videos?
- Do AI-generated articles rank worse in Google Search than human-written ones?
- Do AI-generated articles receive less search traffic than human writing?
- What percentage of workplace communication now contains AI-generated content?
- Does artificial amplification of creator content weaken authentic social proof signals?
- How similar are GPT-generated fake profiles to real human profiles?
- Does detecting AI authorship actually improve social media feed quality?
- How much do humans edit AI-generated text before publishing?
- Do human readers still recognize authors after heavy AI rewriting?