AI researchers debate how close we are to recursive self-improvement
Source: Dwarkesh Patel with John Schulman, Beren Millidge, Charlie O'Neill · 2026-09-11
There’s been a classic thing, almost like Moravec’s paradox, where we think of the AI as, “If it can do this, it’s going to be amazing.” If it can solve these hard maths problems, if it can win at chess, blah, blah, blah... Then it solves these things, and it’s not that impactful. Obviously, it’s somewhat impactful, but not everything.
I agree with that. Humans have a lot of advantages over models now. Each time a new model comes out, it’ll catch up in some of these areas. But you end up getting bottlenecked by the places where the model is weaker and where it has worse judgment, or the models can’t check themselves well enough.
For me, it’s a question of how far off the global optimum of “a learner you could have on a chip” is from the transformer + RL, basically the current recipe. People imagine that once you have an agent which is better than all humans at AI research, even if it’s 0.1% better than all humans, then the fact that you can run hundreds of thousands, if not millions, of these in parallel — and you can run them much faster as chips speed up — is going to outweigh every other bottleneck. You’re eventually going to hit this very fast takeoff with regards to self-improvement.
I think there’s different kinds of research. There’s research in the autoresearch style where the objective is already specified very cleanly and you’re optimizing that objective. I think everyone is picturing that if we continue along this path of making pre-training loss go down and making our environments have the reward on them go up, that’s going to lead to improvement.
But maybe what Ryan is talking about is this much more open-ended type of science which is required for paradigm shifts, where we can’t specify the objective, and the AIs are definitely not able to specify that objective either. We have to be really, really careful about how we specify objectives for any of these things.
I think this is really the key question for any kind of very rapid RSI from current AIs. How well can AIs generalize to learning their own objectives? To have any kind of self-propelling automated loop, we need the AI to propose objectives, optimize them, figure that out, propose a new objective, and have this not go off the rails at any point for a long, long time.
I think distillation is the main thing that fights against the centralizing force. Basically anything that can be learned through RL can be distilled very easily, because it’s a small number of bits. It’s something that you can learn from a small amount of data. If you can get trajectories from the model that show a behavior, you can easily distill it. I think distillation is one of the things that fights centralization.
We’ll probably do some combination of learning from human feedback to absorb the researchers’ taste, and just creating a lot of practice environments which involve doing multi-step research projects. People will in practice do some combination of those two things and, each iteration, patch whatever seems to be most broken in the last iteration. Researchers will be using the AIs a lot and will notice that they have some consistent weaknesses. Those things will either be patched by collecting human feedback or creating environments.
I feel like in AI research especially, it’s very easy to define goals. You could say the loss needs to be 1.3 or something, and no human can get that now. But that’s an extremely measurable, verifiable task. If the AI gets that, then great.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI research automation sustain progress through accelerating feedback loops?- How do AI researchers currently estimate timelines to artificial general intelligence?
- How should superintelligent AI systems be aligned during rapid capability gains?
- Do efficiency gains in AI-assisted development stem from better tools or autonomous improvement?
- Are AI companies already implementing slowdowns in development as claimed?
- Have AI researcher timelines shifted based on recent capability evidence?
- Which bottleneck in the R&D feedback loop is the weakest link today?
- How does automated R&D affect the efficiency of the research process itself?
- What timeline disagreements emerge among researchers about autonomous AI development?
- Does AI research acceleration compound into faster field-wide progress over time?
- How much can computational speed and automation substitute for human scientific judgment?
- Can recursive feedback loops turn AI research automation into genuine progress?
- Which parameters drive the largest uncertainty in AI R&D automation dates?
- Could superhuman research taste accelerate AI development beyond trend extrapolation?
- Can third-party evaluators embedded in labs measure AI-led R&D work reliably?
- What governance approaches do researchers propose for automating AI research?
- Do humans or AI perform better at different research stages?
- How do template requirements limit AI research systems from true autonomy?
- What role should human experts play in AI-driven research ideation loops?
- Why does faster research production force automation of the evaluation process itself?