How many accepted conference papers contain hallucinated citations?
A vendor scan of NeurIPS 2025 papers flagged hundreds of citations that could not be verified online. The question is whether these flags represent genuine hallucinations or unverifiable but real sources, and how many require correction.
GPTZero reports that it scanned 4841 of the 5290 papers accepted to NeurIPS 2025 and found "noticeable traces of AI authorship and hundreds of vibe citations." The company, which sells an AI-text detector and a citation checker called Hallucination Check, frames this as a flaw in the peer review pipeline, not a charge against named organizers or reviewers. Its account of the pressure behind it: NeurIPS submissions rose "from 9,467 to 21,575" between 2020 and 2025, and at a 24.52% main-track acceptance rate each accepted paper "beat out 15,000 other papers despite containing one or more hallucinations." The title's headline count is 100 new hallucinations. The body never gives that number; it says only "hundreds."
The mechanism is a citation check. Hallucination Check uses an in-house agent to flag any citation it cannot find online. GPTZero is explicit that a flag is not a verdict: many archival or unpublished works "can't be matched to an online source," so flags "indicate which sources require further human scrutiny." It defines a vibe citation as one that "likely resulted from the use of generative AI," with errors such as fabricated authors, titles, DOIs or containers, and first names extrapolated from initials. The company claims the tool catches "99 out of 100 flawed citations," with a higher false positive rate because any citation it cannot verify gets flagged. The examples were checked by people: "each of the hallucinations presented here has been verified by a human expert."
Set against the nearest notes, this excerpt is about what the process let through, not how reviewers behaved. Does banning LLM use in peer review change review outcomes? measured reviewer conduct under policy; GPTZero checks the output and its citations, not the reasoning. Can inference scaling help reviewers catch errors humans miss? spends heavy compute per manuscript on proofs and experiments. Citation matching is shallower, but it is mechanical and runs against public records, which is why one pass can cover 4841 papers. The excerpt still routes its flags to people, and Can people reliably spot content made by AI? suggests human judgment is a weak check for AI content on its own, though these flags are checked against online records, not by eye. The vendor-claim problem also recurs: How often do legal AI tools actually hallucinate citations? tested "hallucination-free" marketing independently, while the 99-in-100 figure here is GPTZero's own and untested in the excerpt.
The excerpt does not establish how many flagged citations are confirmed hallucinations. The title's 100 is never reconciled with the body's "hundreds," and the human verification it describes covers "the hallucinations presented here," not the full scan. Its growth figure also needs care: 21,575 over 9,467 is a ratio of about 2.28, or growth of about 128 percent, so "more than 220%" reads as a ratio mislabeled as growth. There is no false positive rate, no per-paper breakdown, no comparison with other venues, and no NeurIPS response. The defensible implication is narrow: unverifiable citations appear often enough in accepted ML papers to justify venue-level checks, and those checks are cheap to run. The scan does not measure how many accepted papers contain fabricated sources. Its figures are the vendor's own until an independent audit reproduces them.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans? How do hallucinated citations emerge in AI scholarly output? What gaps exist between benchmark performance and real deployment outcomes? What are the real-world consequences of AI citation hallucinations? Why do language models hallucinate and how can we prevent it?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
measures reviewer conduct under LLM policies; this excerpt checks what the process let through, not reviewer behavior.
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
a deeper, costlier per-manuscript check of proofs and experiments; citation matching is shallower but mechanical.
-
How often do legal AI tools actually hallucinate citations?
Legal vendors claim their AI research tools eliminate hallucinations, but do they? This preregistered study measures hallucination rates in leading commercial legal-research systems to test those marketing claims.
the same vendor-accuracy problem: an independent test here, a self-reported 99-in-100 catch rate there.
-
Can people reliably spot content made by AI?
This systematic review of 30 studies asks whether human judgment can distinguish AI-generated text, images, and voice from human-created content, and whether detection accuracy has improved as AI becomes more realistic.
explains why flags route to people, and why human judgment alone is a weak fallback.
-
How can conferences detect and handle LLM misuse in peer review?
Explores how ICLR 2026 balanced detection limitations with practical enforcement, distinguishing between acceptable LLM assistance and problematic offloading of reviewing responsibilities.
evidence for: ICLR 2026 kept detector flags as one input because accuracy is imperfect; confirmed fabricated references led to desk rejection
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
- Exploring the use of AI authors and reviewers at Agents4Science
- A Retrospective on the ICLR 2026 Review Process
- Towards End-to-End Automation of AI Research
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
- Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
- AI Hallucination Cases database
- Tortured phrases: A dubious writing style emerging in science. Evidence of critical issues affecting established journals
Original note title
GPTZero finds hundreds of hallucinated citations in accepted NeurIPS 2025 papers — each beat out 15,000 other submissions despite containing hallucinations