Does NeurIPS 2025's LLM disclosure policy match what readers actually need?
NeurIPS 2025 requires LLM disclosure only for methodological use, not writing help. But reader studies suggest disclosure matters more broadly, raising questions about whether the policy's limits adequately protect scientific integrity.
NeurIPS 2025 sorts LLM use by where it lands in the paper. Use that shapes the method has to be described: the policy asks authors to document their methodology and to describe LLM use "in the experimental setup section (or equivalent) if it is an important, original, or non-standard component of the approach." Use only for writing, editing or formatting, which "does not impact the core methodology, scientific rigorousness, or originality," needs no declaration; authors are asked instead to report it in a "statistical analysis survey." Responsibility stays with people whatever the tool. "Only humans are eligible to be authors," and each author is "fully responsible for all the content in your paper, including text, figures, and methodology, regardless of what tools (e.g., LLMs) you have used."
The reasons given are about risk rather than style. Some tools "may retain input data for further model training purposes," which the policy treats as a privacy consideration, and high-level instructions "could potentially result in hallucinations when generating plots, risking scientific integrity." Enforcement is retrospective: NeurIPS "reserves the right to revoke the paper's publication status" for violations, and its example is "using references generated by an LLM without conducting the due diligence to verify correctness, existence and appropriateness." Reviewers get a different and stricter rule. They may not share information about a submission "with anyone or any LLMs," but they may consult LLMs about concepts and phrasing if they take care not to leak submission content.
The ICML 2026 experiment in Does banning LLM use in peer review change review outcomes? tests the reviewer side of this question. Banning LLM use and allowing limited use both barely moved scores, decisions or confidence, and a substantial share of reviewers broke the rules either way. That is evidence on whether stated reviewer rules hold, which the NeurIPS page asserts but does not measure. On the author side, the method-versus-prose line meets the reader evidence in Do readers and writers differ on AI disclosure necessity?. There, readers judged disclosure more necessary than writers did, the judgment rose when AI text was directly adopted, and effort had no significant effect. That points the other way from a writing-only exemption, which covers direct adoption of AI-written text so long as the method is untouched. The policy's documentation demand also lines up with Does iterative prompt engineering undermine scientific validity?, which argues that prompt iteration itself breaks scientific method. NeurIPS asks only that LLM use be described when it is a non-standard part of the approach, a narrower claim.
The excerpt does not establish what happens in practice. It is a policy page. It reports no count of declared or undeclared LLM use, no compliance rate, and no test of whether authors apply the "important, original, or non-standard" test consistently; that judgment is left to the authors. The risks it names (retained inputs, hallucinated plots, unverified references) are stated as concerns, not measured. What follows is narrower than the policy itself: NeurIPS has drawn a clear division of responsibility and a rule for where disclosure belongs, but whether that rule changes what readers and reviewers see is open. A reader who cannot reliably tell LLM-written prose from human prose, as in Can readers tell LLM abstracts from human ones?, cannot see the writing-only use the exemption leaves undeclared. So the policy's line rests on the separate usage survey and on each author's own judgment.
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
tests the reviewer-side rules this policy sets, and finds rule-breaking was common under both ban and limit arms.
-
Do readers and writers differ on AI disclosure necessity?
This vignette study explores whether readers and writers judge the necessity of disclosing AI use differently, and what conditions make disclosure feel more important to each group.
the vignette readers draw the disclosure line at direct adoption of AI text, where this policy draws it at the method.
-
Does iterative prompt engineering undermine scientific validity?
When researchers repeatedly adjust prompts to get desired outputs, does this practice introduce hidden bias and produce unreplicable results? The question matters because LLM-based research is proliferating without clear methodological safeguards.
asks the same documentation of LLM use in methods that this note demands of prompt iteration, without its claim that iteration breaks the method.
-
Can readers tell LLM abstracts from human ones?
Do readers with ML expertise reliably distinguish human-written, LLM-generated, and LLM-edited research abstracts? Understanding this matters for evaluating whether readers can serve as effective gatekeepers against LLM content.
shows readers cannot reliably see LLM writing help, the case the writing-only exemption leaves undeclared.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLM Policy
- ICLR 2026 Response to LLM-Generated Papers and Reviews
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Mapping the Increasing Use of LLMs in Scientific Papers
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
- A Retrospective on the ICLR 2026 Review Process
Original note title
NeurIPS 2025 requires LLM disclosure only where it shapes the method — authors answer for every line and reviewers may not share submissions with LLMs