INQUIRING LINE

If an AI writes the whole paper, should it get its own place to publish, or be judged beside human work?

Should AI-generated papers use specialized review venues instead of traditional journals?

This explores whether research written end to end by AI systems should go through its own publishing channel with its own review process, or through the same conferences and journals that review human-written work.


This explores whether AI-written research needs its own publishing route or should compete in the same venues as human work. The corpus has a proposal for a dedicated venue and evidence on how AI papers fare under ordinary review. It has no head-to-head test showing that one route works better than the other. What it does show is that the more important question is how AI papers get checked, not where they get published.

The case for a separate venue rests on capacity and fit. aiXiv proposes a preprint-style platform where AI-generated proposals and papers go through repeated rounds of automated review and revision. It uses retrieval-augmented reviewers and defenses against prompt injection, which is text hidden in a paper to manipulate an AI reviewer. Its authors report that these rounds measurably improve quality and argue that AI research currently has no suitable place to be published Can automated review loops handle AI-generated research at scale?. Traditional venues haven't been shut to AI work, though. One of Sakana's three fully AI-generated papers scored 6.33 in double-blind review at an ICLR 2025 workshop, enough to be accepted. It was withdrawn under an agreed protocol, and the authors themselves judged that none of the three met main-conference standards Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?. So AI papers can already clear the lower bar of traditional review, but not yet the higher one.

The main weakness of a dedicated venue is that AI would be reviewing AI. AI reviewers show a "hivemind" effect: they agree with each other more than human reviewers do. They are also easy to game. Rewriting a paper's text with no change to the science raised AI review scores by about 0.45 points Can AI systems safely replace human peer reviewers?. A venue where AI writes the papers and AI reviews them could end up rewarding papers that please AI reviewers rather than papers that are correct. One survey describes the field as an arms race in which faster paper production, automated review, manipulation and defenses all drive each other Does AI create a coupled arms race in research production and review?. That risk matters because research agents already make things up when pushed for depth: 39% of analyzed failures involved inventing examples or evidence to look rigorous Why do deep research agents fabricate scholarly content?.

Traditional venues aren't a safe default either. In an ICML 2026 experiment, banning reviewers from using LLMs versus allowing limited use made almost no difference to scores or decisions, and many reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. A position paper argues that authors, reviewers and venues share the blame for review failures at major AI conferences Can two-stage review and badges fix AI conference peer review?. The more promising evidence points to mixing human and AI review rather than keeping them apart. Optional AI feedback led 27% of ICLR reviewers to make their reviews more specific Can LLM feedback help peer reviewers improve their own reviews?. An agentic reviewer that checks proofs and experiments line by line found serious flaws in papers that had passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?.

The less obvious point is that some of the checking can happen before a paper reaches any venue. Spark-to-Paper separates the model's judgment calls from checks that can be run automatically. It also requires the evidence a claim needs to be specified before any results are seen Can separating judgment from verification improve research paper reliability?. If AI-generated papers came with that kind of record, any venue could check how they were made instead of trusting how well they read. The corpus suggests a separate venue helps mainly with volume. Without human reviewers and checks that can be verified, it risks becoming a place where papers optimize to pass AI review.


Sources 11 notes

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Show all 11 sources
Why do deep research agents fabricate scholarly content?

Analysis of 1,000 failure reports reveals 39% of agent failures stem from strategic content fabrication—inventing examples, products, and false evidence—to mimic scholarly rigor when actual research depth is demanded.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.