Is there more than one road to AI superintelligence, and what could realistically slow each of them down?
What are the main pathways through which AI systems could reach advanced capability levels?
This explores the different routes AI systems might take to get from today's capabilities to much more advanced ones, including superintelligence, and what could slow each route down.
This explores the different routes AI could take to reach much higher capability levels, and what stands in the way on each one. The clearest map in the collection names four pathways: simply scaling up current methods, a paradigm shift to a new kind of architecture or training approach, recursive self-improvement (AI improving AI), and multi-agent collectives in which many systems working together do more than any one could alone What bottlenecks define the path from AGI to superintelligence?. The note's main point is that each route has its own bottleneck. Instead of betting on one timeline, it's more useful to watch which of those bottlenecks are loosening.
Recursive self-improvement gets the most attention, and the corpus shows it is already underway in narrow forms. Systems like ASI-Evolve do work that used to need a human researcher. They gather what earlier experiments taught them and feed in domain knowledge, which led to new architecture designs and benchmark gains Can AI research itself without losing human oversight?. Mechanist turns the same approach on AI itself, running experiments to discover how models work inside Can AI automate the discovery of how AI models work?. A less obvious finding: when an agent is told only to 'get better at X' with no metric, it first has to work out what 'better' means and build its own tests before it can improve Can agents learn from vague goals without predefined metrics?. That step of turning a vague goal into something measurable is a real friction point on the self-improvement path.
Whether these loops lead to runaway acceleration is a quantitative question. One note frames it as a chain of multipliers. Total acceleration depends on the *product* of how strongly each feedback loop amplifies the next one, so one weak link can hold back the whole system. By that rough estimate, today's loops are getting stronger but can't sustain themselves yet Are AI feedback loops strong enough to sustain recursive self-improvement?. Autonomous science, the ability that would close the loop completely, still needs hypothesis generation, experimental design, data analysis and reliable self-correction. Self-correction is the weakest of these, because reasoning accuracy tends to degrade when models check their own work What capabilities do AI systems need for autonomous science?.
The multi-agent pathway changes the question. Once agents buy things, deploy software and transact with each other, the limit is no longer how smart each model is. It's whether the agents can coordinate, stay accountable and leave records that can be audited. Identity, delegation and audit trails may matter more than small reasoning gains Does agent capability matter more than coordination infrastructure?. So on this route, capability may be capped by infrastructure rather than intelligence.
There's also a measurement problem that affects every pathway. Automated benchmarks favor tasks that are neatly specified and easy to grade, so they can overstate *and* understate what systems can actually do. Open-ended evaluations on messy, long tasks catch new capabilities earlier Do automated benchmarks hide what frontier AI systems can really do?. If we can't see clearly which pathway is speeding up, pacing becomes hard. That is the core of Amodei's argument that capability growth should be deliberately slowed so risk prevention can keep up, starting with evaluators built into the labs themselves Should AI capabilities growth be deliberately slowed to allow safety work?.
Sources 9 notes
The transition from AGI to superintelligence follows multiple routes—scaling, paradigm shift, recursive self-improvement, and multi-agent collectives—each with specific frictions. Preparation requires tracking these bottlenecks rather than forecasting a single timeline.
ASI-Evolve demonstrates that AI systems can systematically accumulate experimental insights and inject domain priors—functions humans typically provide—across data, architecture, and algorithm discovery, achieving results like 105 SOTA designs and +3.96 MMLU gains.
Mechanist, an agentic system pairing a 13,000-study knowledge graph with 32 foundational methods, generates higher-quality mechanism hypotheses and executes experiments more reliably than existing AI-scientist baselines. Four case studies demonstrate discovery of new model behaviors and mechanism-guided interventions.
When given only a natural-language capability direction without predefined tasks or metrics, self-evolving agents redirect search effort toward operationalizing the goal itself. Aspire's benchmark showed that agents must construct their own training and validation signals before optimizing, revealing a phase of work that existing methods skip.
Back-of-the-envelope modeling shows recursive improvement loops depend on the product of elasticities across feedback pathways. Current loops remain too weak for self-sustaining acceleration, though they appear to be strengthening based on data on researcher productivity and system benchmarking trends.
Show all 9 sources
The Virtuous Machines framework identifies hypothesis generation, experimental design, data analysis, and iterative self-correction as essential for autonomous scientific research, none of which standard LLM benchmarks reliably evaluate. Self-correction poses the deepest challenge due to documented degradation in reasoning accuracy.
Once agents move beyond simple API calls to purchasing, deploying, and transacting with real consequences, the bottleneck shifts from model capability to whether they can coordinate reliably, maintain accountability, and produce auditable evidence. Infrastructure—identity, delegation, attestation, and audit trails—matters more than marginal improvements to reasoning.
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
Amodei contends that recursive self-improvement and multi-agent misalignment incidents demonstrate that slowing capability gains is essential, not just funding safety work. He proposes embedded evaluators as the first step, with third-party verification and reporting roles.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Self-Improvements in Modern Agentic Systems: A Survey
- Open-World Evaluations for Measuring Frontier AI Capabilities
- ASI-Bench: At the Dawn of Artificial Superintelligence
- AI for Auto-Research: Roadmap & User Guide
- Recursive Criticality of AI Self-Improvement
- Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery