INQUIRING LINE

Can plain code, not another AI's opinion, stop an automated research system from publishing claims that don't hold up?

What deterministic checks prevent AI research systems from publishing unsound claims?

This explores which mechanical, rule-based safeguards (as opposed to asking another AI to judge) can stop automated research pipelines from producing papers whose claims don't hold up.


This explores which mechanical, rule-based safeguards, as opposed to asking another AI for its opinion, can stop automated research pipelines from publishing claims that don't hold up. The corpus gives a fairly specific answer: leave the model's judgment in place, but surround it with checks that the model has no say over. The clearest version is the Spark-to-Paper design, which splits paper writing into separate skills. The parts that need judgment stay with the model. Everything that can be run and verified (computing a statistic, checking that a figure matches the data, confirming that a cited result exists) becomes executable code. The key move is that the system has to state what evidence would support a claim before it sees the results. That way, the paper's consistency no longer depends on the model being right Can separating judgment from verification improve research paper reliability?.

That 'commit before you look' rule matters because the alternative has already been shown to run at industrial scale. One demonstration produced 288 finance papers from 96 statistically significant signals. Each paper had a theory invented after the fact and fabricated citations. This is HARKing (hypothesizing after the results are known), automated Can AI generate hundreds of fake academic papers automatically?. A broader survey finds the same pattern across the research lifecycle: AI produces plausible outputs faster than anything can verify them. In agentic research, 39% of failures come from fabricated content and 32% from retrieval failures, not from misunderstanding Can AI verify research outputs as fast as it generates them?. Most of those failure types are the kind a mechanical check can catch: does this citation resolve, does this number come from this run, was this hypothesis registered before the data came in?

A second group of safeguards protects whatever AI judging remains. One note lists four mechanical moves. Run the unarguable checks first, before any contestable ones. Measure the judge's accuracy against human labels. Hide test data from the system proposing answers. Plant known-bad cases as alarms that should always trip Can deterministic checks protect LLM judges from failure?. None of these needs the LLM to vouch for itself, and that matters because LLM judges can be gamed without any access to the model. Fake references and polished formatting raise scores whatever the content says Can LLM judges be tricked without accessing their internals?. An unsound paper decorated with citations is exactly what an unguarded AI reviewer would reward.

The less obvious lesson is that the strongest checks so far look less like rules and more like doing the work again. PAT, an agentic reviewer that spends extra compute re-checking proofs and experiments line by line, found serious flaws in STOC and ICML papers that human reviewers had passed Can inference scaling help reviewers catch errors humans miss?. Agent-based judges that collect evidence cut judge drift from 31% to 0.27%, but their memory module passed errors along from step to step. Checkers need their own error isolation Can agents evaluate AI outputs more reliably than language models?. The Darwin Gödel Machine dropped formal proof in favor of running benchmarks empirically. Running the code is often the only deterministic test available Can AI systems improve themselves through trial and error?.

The corpus has a gap here: no note shows a full set of checks that would reliably stop a bad automated paper end to end. Tools for keeping errors visible and recoverable exist only in fragments How can we measure whether AI errors stay visible and recoverable?. The current backstop is still human: AI Scientist-v2's authors withdrew their accepted workshop paper themselves Can AI systems generate research papers that pass peer review?. One philosophical note argues that AI output resembles hearsay, which traditional verification tools like citation and peer review were never designed to process Does AI-generated knowledge have the same structure as hearsay?. Read that way, deterministic checks are an attempt to rebuild the evidence trail that AI output removes.


Sources 11 notes

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Can AI generate hundreds of fake academic papers automatically?

A demonstration showed LLMs generating 288 complete finance papers from 96 statistically significant signals, each with invented theoretical justifications and fabricated citations, proving academic HARKing can be automated at scale.

Can AI verify research outputs as fast as it generates them?

AI can produce plausible research outputs faster than it can prove them correct or meaningful, shifting the bottleneck from authorship to verification. Evidence shows 39% of agentic research failures stem from content fabrication and 32% from retrieval failures, not comprehension—and the gap widens precisely where novelty and scientific judgment matter most.

Can deterministic checks protect LLM judges from failure?

Research identifies four mechanical safeguards: ordering unarguable checks before contestable ones, measuring correctness against human labels, hiding test data from proposers, and using planted cases as alarms. None requires the LLM itself to verify compliance.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Show all 11 sources
Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

How can we measure whether AI errors stay visible and recoverable?

Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Does AI-generated knowledge have the same structure as hearsay?

AI output shares all defining features of hearsay: testimony at remove, modification in retelling, unattributable origin, and unverifiability against stable sources. This means Enlightenment verification tools—citation, archiving, peer review, evidentiary chains—cannot process AI output by design.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.