ASI-Bench: At the Dawn of Artificial Superintelligence
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today’s AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems’ capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation.
Introduction. A central challenge on the path toward artificial superintelligence (ASI) is whether AI can move beyond mastering existing human knowledge to explore unfamiliar problems, develop new solutions, and turn them into verifiable results. Today’s AI systems derive much of their capability from learning, compressing, and applying the accumulated knowledge of humanity, and have made rapid progress in scientific reasoning, coding, data analysis, and agentic execution [1, 2, 3, 4, 5, 6, 7]. Yet existing evaluations largely test these capabilities either through problems with known answers or through tasks whose methods and procedures are substantially specified by humans. They therefore provide limited evidence about whether AI can autonomously conduct scientific research when both the problem and the path to a solution are open-ended. In this work, we ask a more direct question: how far can current AI systems independently explore and execute project-level scientific research as human methodological guidance is progressively withdrawn?
Discussion / Conclusion. Beyond a single leaderboard, ASI-Bench is intended to serve as shared research infrastructure for the scientific and AI communities. Its B1–B4 structure and domain-level analyses enable controlled comparisons across models and agents, while revealing where progress occurs and where human methodological guidance remains necessary. Yet no fixed benchmark—and no single research team—can fully represent the breadth of difficult, meaningful, and verifiable problems that future AI systems must confront. The continued value of ASI-Bench therefore depends on collective participation from researchers working across disciplines and at the frontier of model development.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How should human oversight be integrated with autonomous AI systems?- Where do human researchers retain competitive advantage over autoresearch systems?
- Where is human judgment still essential in AI-assisted research?
- How should safeguards be built into AI research pipelines?
- Which research stages are actually high-leverage decision points for human intervention?
- Why do major AI breakthroughs require human-discovered data and method combinations?
- Which research collaboration skills should AI systems develop first?
- What tasks do users actually want AI to handle versus what can it automate?
- Which task characteristics determine whether AI can displace them first?
- How should domain-specific AI be evaluated differently from general benchmarks?
- Why do benchmark scores not capture the true nature of AI systems?
- How should single-axis benchmarks account for separable capability dimensions?
- Can the scaling law for discovery extend beyond architectures to agentic systems?
- Does architectural discovery follow an empirical scaling law like neural networks?
- Do autonomous architecture discoveries follow predictable scaling laws like human research?
- What scaling laws govern autonomous architecture discovery in AI systems?