Do researchers who write with AI actually end up citing a wider and more recent range of sources?
Do LLM adopters actually cite more diverse and younger research?
This explores whether researchers who use LLMs to write papers end up citing a broader range of work and more recent work, and whether this collection has evidence on that.
This explores whether researchers who use LLMs in their writing cite a wider and more recent range of work. The direct answer: this collection doesn't have a study that measures this. None of the retrieved notes compare how old or how varied the reference lists are for LLM adopters versus non-adopters. What the corpus does have is several nearby findings, and together they suggest why the question is harder to answer than it looks.
Start with what citations mean to readers. In an analysis of 24,000 search-assistant interactions, users rewarded irrelevant citations almost as much as relevant ones, so the number of citations works as a trust signal that has little to do with whether they support anything Do users trust citations more when there are simply more of them?. That matters here. If LLM-assisted writing produced longer or more varied reference lists, readers and reviewers would probably reward them whether or not those references did real work. A more diverse bibliography and a better-grounded paper are not the same thing. The worst version of this has already reached enforcement: ICLR 2026 desk-rejected papers with confirmed fabricated references. Program chairs treated hallucinated citations as the one LLM problem clear-cut enough to act on, while detector flags went to human area chairs as one input among several How can conferences detect and handle LLM misuse in peer review?.
Whether LLMs widen what researchers draw on is an open question, and the corpus points both ways. On the side of breadth, LLM-generated research ideas were rated more novel than expert ideas, apparently because the models combine concepts that expertise would rule out Do language models generate more novel research ideas than experts?. That novelty dropped once the ideas were actually carried out Why do LLMs generate more novel research ideas than experts?. On the side of narrowness, a study of 1,006 LLM papers found that their engagement with psychology runs through a few familiar routes (CBT, stigma theory, the DSM) while whole traditions go largely uncited Why do AI researchers cite only narrow psychology pathways?. That study doesn't separate LLM-assisted authors from others, but it shows the field's citation habits can be quite narrow whatever tools authors use.
There's also a quieter reason to doubt that LLMs help authors cite better. Language models handle text but not the social context that gives an expert's claim its weight: reputation, track record, standing in the field Can language models distinguish expert arguments from common assumptions?. A model suggesting references can't easily tell a foundational paper from a widely repeated assumption. It might pull in a more varied set without improving how the sources are chosen. Finally, a warning from peer review: across 125,000+ reviews, the apparent bias of LLM-assisted reviewers toward LLM-assisted papers disappeared once paper quality was controlled for Do LLM reviewers actually favor LLM-written papers?. Any finding that LLM adopters cite differently would need the same check, because adopters may differ from non-adopters in seniority, subfield, or paper quality in ways that alone shape a bibliography.
So the honest takeaway is this. The collection can't tell you whether adopters cite more diverse or more recent work. It can tell you that the number of citations is a weak sign of quality, that fabricated references are the most serious documented risk, and that any reported pattern in adopters' reference lists is likely to reflect who the adopters are rather than what the tool does, unless the study controls for that.
Sources 7 notes
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.
Research shows LLM-generated ideas are statistically more novel than expert-produced ideas, but LLMs struggle to evaluate quality—automated evaluation overestimates by 60%. When executed, LLM ideas drop significantly on all metrics, suggesting novelty without feasibility.
Analysis of 1,006 LLM papers shows CBT, stigma theory, and DSM dominate mental health citations while developmental neuropsych and psycholinguistics remain underused. This narrow foundation risks building AI tools on incomplete psychological understanding.
Show all 7 sources
LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- A Retrospective on the ICLR 2026 Review Process
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Stop Automating Peer Review Without Rigorous Evaluation