Scientific papers all come packed as one fixed PDF bundle — what if AI let each piece move separately?
How does document form shape what kinds of evidence social science can present?
This explores how the standard container for social science, the peer-reviewed PDF, limits what counts as evidence and how it gets presented, and what might change if AI loosened that container.
This explores how the format social scientists publish in, mostly the peer-reviewed PDF, shapes what evidence they can show and how readers weigh it. The corpus has one note that takes this on directly, and several others that come at it from the side. The direct one is Munger's argument that the PDF is a bundle Can AI help social science move beyond the peer-reviewed PDF?. A single paper has to act as an archive, a literature review, a theory statement, a methods record and a results report all at once. Evidence only enters the record if it fits into that fixed sequence. Munger's surprising claim is that the main obstacle to change isn't AI capability. It's the document form itself and the career incentives built around it. The tools to split these functions into separate, recombinable pieces already exist. The PDF is what keeps them stuck together.
You can see why this matters when you look at the new kinds of evidence AI is producing that don't fit a results section. LLM-driven simulations guided by causal models can test social hypotheses in settings like negotiation, bail and job interviews. They reliably show which direction an effect goes, but they can't say how big it is Can structural causal models automate social science with language models?. A paper format built around effect sizes and significance tests has no natural place for evidence that is directional only. A related idea is that LLM outputs shouldn't be treated as observations at all. They are draws from the model's learned assumptions, a 'subjective prior' shaped partly by how the prompt was written. They should influence conclusions only through an explicitly stated trust weight Should we treat LLM outputs as real empirical data?. That kind of evidence comes with a dial attached, and a static PDF has nowhere to put the dial.
Form also shapes credibility separately from content. An arXiv preprint shaped public debate about AI and science well before anyone reviewed it. MIT's later withdrawal request couldn't undo that Can unreviewed preprints shape scientific debate before peer review?. Looking like a paper was enough to do the work of evidence. At the reader's end, people prefer answers with more citations, and irrelevant citations raise trust almost as much as relevant ones Do users trust citations more when there are simply more of them?. The visible apparatus of scholarship works as a trust signal on its own. Research on persuasion points the same way: claims slipped in as assumed background persuade more than claims stated outright Why are presuppositions more persuasive than direct assertions?. That's worth keeping in mind for a genre where the literature review sets out what readers are expected to take as settled.
Some notes suggest what an unbundled form could look like. One paper-generation system requires authors to specify their evidence before seeing results, and keeps model judgment separate from checks a machine can verify Can separating judgment from verification improve research paper reliability?. That builds something like pre-registration into how the document is made. On the reading side, retrieval works better when it first builds a map of what role each part of a document plays Can building a document map first improve retrieval over long texts?. This hints that a document's internal roles, like 'this is the method' and 'this is the finding', could become handles that machines and readers use directly. There's also a caution. Horning argues that answer-style AI output presents frozen facts while hiding the social processes that produced them Do LLMs obscure the historical processes behind their answers?. Replacing the PDF with AI-generated summaries could hide even more of how the evidence was made, not less.
In short, the corpus has one sharp argument (Munger's) and a set of adjacent findings. Together they suggest that the form of a document decides which evidence can be shown, how much it gets trusted, and how easily it can be checked. The corpus doesn't yet include empirical studies comparing what social scientists actually report under different publication formats, so that part of the question stays open.
Sources 9 notes
Munger contends that the peer-reviewed PDF combines distinct functions—archive, literature review, theory, methods, results—that AI could separate into recombinable forms. He identifies document form and academic incentives, not AI capability limits, as the bottleneck preventing this transition.
LLMs guided by structural causal models can propose and test causal hypotheses across negotiation, bail, interview, and auction scenarios. Simulations reveal effect directions reliably but not magnitudes, making them useful for directional social science.
Foundation Priors framework shows that LLM-generated text reflects the model's learned patterns and user's prompt choices, not ground truth. Such outputs should only influence inference through explicitly parameterized trust weights, not be treated as equivalent to real evidence.
MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Show all 9 sources
Experimental evidence shows presuppositions with additive, iterative, and factive triggers persuade audiences more than assertions, especially for discourse-new content. The mechanism: presuppositions bypass evaluative scrutiny by presenting claims as already-accepted background.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
MiA-RAG inverts standard RAG by summarizing documents first, then conditioning retrieval on that global view. This approach recovers discourse structure that bag-of-chunks retrieval destroys, making scattered evidence findable by their document role rather than surface similarity alone.
Horning argues LLMs exemplify Lukács's concept of petrified factuality by offering static facts while obscuring the dynamic social relations that produced them. He treats this as a deliberate social function of the technology, supported by evidence that lower AI literacy correlates with greater receptivity.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Towards Automating Scientific Review with Google's Paper Assistant Tool
- A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Mapping the Emerging Social Science of Large Language Models
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- Automated Social Science: Language Models as Scientist and Subjects