How much guidance do AI systems need to conduct research independently?
ASI-Bench tests whether AI can explore open-ended research problems by progressively removing human methodological guidance. This matters because existing benchmarks cannot distinguish between AI that follows instructions well and AI that can autonomously discover and verify new knowledge.
ASI-Bench makes a design claim about what a research benchmark has to vary. The abstract argues that artificial superintelligence requires AI to move from mastering existing knowledge toward "exploring the unknown, creating new knowledge, and turning new ideas into verifiable results," while existing benchmarks "primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance." Neither shows whether AI can conduct research when "both the problem and the path to a solution are open-ended." The authors present ASI-Bench as the first benchmark to jointly evaluate innovative exploration and autonomous scientific execution, and the first to "progressively withdraw human methodological guidance within the same research project."
The mechanism is the withdrawal itself. The benchmark has 60 project-level tasks across 11 scientific domains, and it "progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results." The design implies that a change in performance between levels can be read as an effect of the removed guidance, since the project stays the same. The discussion points to a "B1–B4 structure" and domain-level analyses that enable "controlled comparisons across models and agents, while revealing where progress occurs and where human methodological guidance remains necessary." The question becomes a curve (how far capability extends as guidance recedes) instead of a single pass-or-fail score. The paper backs the tasks with over 40 experts, 31,000+ human hours, expert review, AI-assisted auditing, sandbox execution, and scorer validation.
This sits beside Do automated benchmarks hide what frontier AI systems can really do? as a different answer to the same complaint. That note says benchmarks favor precisely specified, auto-gradable tasks and turns to small-sample qualitative analysis of messy real work. ASI-Bench keeps the benchmark form, with sandboxes and validated scorers, and makes the level of specification the experimental variable. It also offers a second axis for testing Where does AI assistance become unreliable in research?. That note locates the boundary by lifecycle stage, and a guidance-withdrawal design would locate it by how much methodology the human supplies. Its pairing of exploration with execution roughly parallels What capabilities do AI systems need for autonomous science?, though the excerpt does not say how the two line up.
The excerpt reports no results. It gives no scores, no model comparisons, no definition of what B1 through B4 mean, no account of how guidance is operationalized at each level, and no description of how "innovative" exploration is judged. It cannot say whether capability degrades gradually or abruptly as guidance recedes; the introduction poses that as its question and leaves it open. The discussion also concedes that "no fixed benchmark" can represent the full breadth of meaningful, verifiable problems. What follows at this strength is that the guidance-withdrawal design is a well-motivated way to measure autonomy, and any claim about where human methodological guidance stays necessary, which bears on Can human-AI research teams improve faster than autonomous AI systems?, has to wait for the results.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can brute-force automated research substitute for iterative depth and human research intuition? When should work require human-AI partnership versus full automation?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do automated benchmarks hide what frontier AI systems can really do?
Benchmarks optimize for auto-gradable, short, cheap tasks. But real AI capability emerges in long-horizon, messy, open-ended work. How much capability are we missing—or wrongly inflating—by relying on benchmark scores alone?
contrasts: both critique benchmarks, but ASI-Bench keeps scored tasks and varies guidance instead of moving to qualitative log analysis
-
Where does AI assistance become unreliable in research?
This explores whether AI capability follows a sharp boundary in research tasks, and what determines which side of that line a task falls on. Understanding this matters because it reveals where humans must stay in control.
guidance withdrawal could locate the autonomy boundary along a second axis besides lifecycle stage
-
What capabilities do AI systems need for autonomous science?
Explores whether current AI benchmarks actually measure what's required for independent scientific research—hypothesis generation, experimental design, data analysis, and self-correction—or if they test only adjacent skills.
same gap in existing benchmarks; ASI-Bench pairs exploration with execution
-
Can human-AI research teams improve faster than autonomous AI systems?
Explores whether keeping humans actively involved in AI research collaboration accelerates paradigm discovery compared to fully autonomous self-improvement, and what safety advantages this preserves.
a guidance-withdrawal curve would test where human input is still needed
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- ASI-Bench: At the Dawn of Artificial Superintelligence
- Atria Dawn: The Dawn of Agentic Superintelligence
- AI for Auto-Research: Roadmap & User Guide
- ASI-Evolve: AI Accelerates AI
- AI-Researcher: Autonomous Scientific Innovation
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Interactive Evaluation Requires a Design Science
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Original note title
progressively withdrawing methodological guidance within the same research project measures how far AI can proceed on its own