SYNTHESIS NOTE
Topics›Evaluations›this note

How much guidance do AI systems need to conduct research independently?

ASI-Bench tests whether AI can explore open-ended research problems by progressively removing human methodological guidance. This matters because existing benchmarks cannot distinguish between AI that follows instructions well and AI that can autonomously discover and verify new knowledge.

Synthesis note · 2026-09-25 · sourced from Evaluations

ASI-Bench makes a design claim about what a research benchmark has to vary. The abstract argues that artificial superintelligence requires AI to move from mastering existing knowledge toward "exploring the unknown, creating new knowledge, and turning new ideas into verifiable results," while existing benchmarks "primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance." Neither shows whether AI can conduct research when "both the problem and the path to a solution are open-ended." The authors present ASI-Bench as the first benchmark to jointly evaluate innovative exploration and autonomous scientific execution, and the first to "progressively withdraw human methodological guidance within the same research project."

The mechanism is the withdrawal itself. The benchmark has 60 project-level tasks across 11 scientific domains, and it "progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results." The design implies that a change in performance between levels can be read as an effect of the removed guidance, since the project stays the same. The discussion points to a "B1–B4 structure" and domain-level analyses that enable "controlled comparisons across models and agents, while revealing where progress occurs and where human methodological guidance remains necessary." The question becomes a curve (how far capability extends as guidance recedes) instead of a single pass-or-fail score. The paper backs the tasks with over 40 experts, 31,000+ human hours, expert review, AI-assisted auditing, sandbox execution, and scorer validation.

This sits beside Do automated benchmarks hide what frontier AI systems can really do? as a different answer to the same complaint. That note says benchmarks favor precisely specified, auto-gradable tasks and turns to small-sample qualitative analysis of messy real work. ASI-Bench keeps the benchmark form, with sandboxes and validated scorers, and makes the level of specification the experimental variable. It also offers a second axis for testing Where does AI assistance become unreliable in research?. That note locates the boundary by lifecycle stage, and a guidance-withdrawal design would locate it by how much methodology the human supplies. Its pairing of exploration with execution roughly parallels What capabilities do AI systems need for autonomous science?, though the excerpt does not say how the two line up.

The excerpt reports no results. It gives no scores, no model comparisons, no definition of what B1 through B4 mean, no account of how guidance is operationalized at each level, and no description of how "innovative" exploration is judged. It cannot say whether capability degrades gradually or abruptly as guidance recedes; the introduction poses that as its question and leaves it open. The discussion also concedes that "no fixed benchmark" can represent the full breadth of meaningful, verifiable problems. What follows at this strength is that the guidance-withdrawal design is a well-motivated way to measure autonomy, and any claim about where human methodological guidance stays necessary, which bears on Can human-AI research teams improve faster than autonomous AI systems?, has to wait for the results.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can brute-force automated research substitute for iterative depth and human research intuition? When should work require human-AI partnership versus full automation?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 98 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

progressively withdrawing methodological guidance within the same research project measures how far AI can proceed on its own