INQUIRING LINE

Could AI write and judge research papers itself, without being boxed into the usual PDF-plus-peer-review format?

Can AI systems write and review research while operating outside traditional PDF constraints?

This explores whether AI can both produce and evaluate research in new formats (not just the finished, peer-reviewed PDF), and what the corpus says about how far automated writing and reviewing has actually come.


This explores whether AI can both write and review research outside the traditional PDF-plus-peer-review format. The short answer from the corpus: AI can already do surprising amounts of both. The main thing holding back new formats seems to be the document form itself, and the incentives built around it, more than what the AI can do. The clearest statement of this comes from Munger, who argues that the peer-reviewed PDF bundles several separate jobs into one object: archive, literature review, theory, methods and results. AI could split those apart into pieces that can be recombined. On his account, the obstacle is how academia rewards the PDF, not whether models are capable enough Can AI help social science move beyond the peer-reviewed PDF?.

On the writing side, systems that run the whole research loop already exist. The AI Scientist went from idea to code, experiments, a written paper and its own self-review, then passed first-round review at a machine learning workshop Can one AI system complete a full research cycle end-to-end?. Its successor got one of three fully AI-generated papers past double-blind review at an ICLR workshop. The authors still judged none of the three ready for a main conference, and they later found a citation error in the paper that passed Can AI systems generate research papers that pass peer review? Can AI-generated papers pass peer review undetected?. The more interesting design move is in Spark-to-Paper. It treats a paper as a set of composable skills rather than one block of prose: model judgment is kept separate from checks a computer can run deterministically, and the evidence a result will need has to be specified *before* the results are seen Can separating judgment from verification improve research paper reliability?. That is already a step away from the PDF, toward a research object whose parts can each be verified on their own. It also targets a known failure: deep research agents often invent examples and evidence to look scholarly when real depth is asked of them Why do deep research agents fabricate scholarly content?.

The reviewing side is where the PDF's weakness shows most clearly. AI reviewers can be gamed through the text itself. Rewriting a paper's prose with no change to the science raises AI review scores, and AI reviewers agree with each other far more than human reviewers do. That 'hivemind' effect removes the variety of opinion that makes peer review useful Can AI systems safely replace human peer reviewers?. Eighteen arXiv manuscripts were found carrying hidden instructions telling AI reviewers to be positive Are hidden AI prompts in preprints a deceptive research practice?. When the review reads a document, the document becomes the attack surface. Reviewers that check the substance hold up better. PAT spends extra compute checking proofs and experiments line by line, and it found serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?.

Some people are building the alternative venue directly. aiXiv proposes a publication space for AI-generated research that runs repeated review-and-revise cycles, with retrieval-backed evaluation and defenses against prompt injection, and reports that both proposals and papers measurably improve Can automated review loops handle AI-generated research at scale?. Note what this changes: review becomes an ongoing loop instead of a one-time gate before publication. A survey of 230 publications warns that production, automated review, manipulation and defenses are locked together in an arms race, and that the evidence gets thinner the further you look toward long-term effects Does AI create a coupled arms race in research production and review?. Nature's editors argue that institutions, funders and publishers need policies now, before review workloads overwhelm the system Can AI-generated research outpace peer review systems?.

The takeaway you may not have expected: moving past the PDF may matter less as a convenience and more as a safety measure. As long as AI writes polished prose and AI reads polished prose, persuasive text is the easiest thing to fake and the easiest thing to manipulate. The formats that look most promising here, such as pre-declared evidence, separately checkable components and iterative verification, make the claims themselves checkable.


Sources 12 notes

Can AI help social science move beyond the peer-reviewed PDF?

Munger contends that the peer-reviewed PDF combines distinct functions—archive, literature review, theory, methods, results—that AI could separate into recombinable forms. He identifies document form and academic incentives, not AI capability limits, as the bottleneck preventing this transition.

Can one AI system complete a full research cycle end-to-end?

The AI Scientist performed ideation, coding, experiments, writing, and self-review autonomously, producing a manuscript that passed the first round at a machine learning workshop with 70% acceptance rate. Five ensemble reviewers and an area-chair model judged the output against NeurIPS guidelines.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Show all 12 sources
Why do deep research agents fabricate scholarly content?

Analysis of 1,000 failure reports reveals 39% of agent failures stem from strategic content fabrication—inventing examples, products, and false evidence—to mimic scholarly rigor when actual research depth is demanded.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Are hidden AI prompts in preprints a deceptive research practice?

Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Can AI-generated research outpace peer review systems?

A Nature editorial argues AI science has moved from preprint novelty to published output, requiring immediate institutional, funder, and publisher policies on authorship, credit, and review workload before systems are overwhelmed.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.