SYNTHESIS NOTE
Topics›Domain Specialization›this note

Does more thinking time improve AI-generated research hypotheses?

Co-Scientist's builders tested whether giving a multi-agent AI system more compute during hypothesis generation produces better scientific ideas. They measured Elo ratings across 203 research goals to explore this scaling relationship.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The paper's central claim is that Co-Scientist, a multi-agent system built on Gemini, produces better research hypotheses the more test-time compute it is given. The builders measured Elo ratings of hypotheses and proposals "over the course of its thinking and computation (i.e. the tournament of hypotheses)" across 203 research goals entered until February 3, 2025, and report "improving hypothesis quality over time." Three biomedical cases carry the validation: repurposing candidates and synergistic combinations for acute myeloid leukemia tested in vitro, novel epigenetic targets for liver fibrosis tested in human hepatic organoids, and a mechanism of mobile genetic element transfer that the group had found but not yet published, which the system "independently recapitulated." These results are the builders' own, and the excerpt reports no outside replication.

The mechanism is the loop the authors call "generate, debate, evolve." Generation agents propose hypotheses under plausibility, novelty, testability and safety criteria. A Reflection agent with search tools is credited with preventing "seemingly novel but implausible hypotheses," and a Ranking agent runs a scientific debate that the ablation study says "significantly improved the ranking of hypotheses and reduced positional bias." A tournament compares hypotheses through win and loss patterns, and an Evolution agent refines the leaders. The authors frame the result as scaling "research ideation with test-time compute rather than exhaustive generation." The ablations show which components matter; the Elo curve is the scaling evidence. The expert comparison is described only as "small-scale," and the automated evaluation over 15 expert-curated goals is also the builders' own measure.

The closest neighbor is How do collaboration rules shape hypothesis quality?, which also runs hypothesis generation as a population search with explicit selection and replacement. HypoEvolve's contribution is to name each collaboration rule as a testable step; the Co-Scientist excerpt reports the tournament outcome without naming the rules that could be varied. Can AI research itself without losing human oversight? also closes a generate-and-evaluate loop, but it distills outcomes into reusable insights where Co-Scientist ranks hypotheses by tournament and passes a context memory forward. Can decentralized teams outperform central planners in long-running science? argues for decentralized teams; Co-Scientist's expert supervision and Meta-review agent are a different coordination choice. The sharpest contrast is with Can human-AI research teams improve faster than autonomous AI systems?. The paper speaks of a "self-improving loop" and "recursive self-improvement with increasing compute," the vocabulary that note treats with caution, yet it keeps scientists in the loop for supervision and feedback, so the excerpt sits between the two positions.

The excerpt does not establish that the Elo gains translate into faster discovery. The ratings come from the system's own tournament, and the excerpt does not say how those judgments were checked against outside raters. The wet-lab work covers three cases, which the authors themselves call "preliminary," and they name further limits: reliance on open-access literature that omits paywalled prior art and negative results, hypotheses that inherit "the mixed and contradictory quality of the source literature," and inherited hallucination. What the excerpt supports is narrower: on the builders' own measure, more compute raised tournament ratings, and three biomedical cases produced experimentally tested candidates. The scaling curve is evidence that the tournament works as designed, not evidence about discovery rates. The wet-lab cases are the stronger check, and there are only three of them.

Inquiring lines that read this note 9

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI-assisted research sacrifice exploration breadth for productivity gains? What human oversight must AI research systems have? Can AI systems perform peer review as effectively as humans? Why do LLM research ideation systems generate novelty but lack diversity?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 92 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Co-Scientist's builders report that hypothesis quality rises with test-time compute in a generate-debate-evolve tournament