Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›What factors determine agentic sys…›this line of inquiry
What fundamental constraints limit how effectively agents can improve themselves?
A broader line of inquiry — a family of 45 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 45
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do evolutionary archives let agents improve themselves without formal proof?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- Can agents design their own objective functions as part of learning?
- Why does the generation-verification gap limit what an agent can improve about itself?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- Can agents improve reliably without an external standard?
- What makes a self-improvement win untrustworthy and why hide evaluations from agents?
- Can agents learn to compress verified evidence and unresolved constraints into a compact improvement state?
- Does co-evolution empirically outperform single-entity self-improvement in standard evaluations?
- Should we train the evolver or the executor when building self-improving agents?
- How do hidden evaluations and out-of-distribution benchmarks address recursive self-improvement risks?
- Can a progressively stricter evaluator act like a curriculum for improving agents?
- What makes an evaluation criterion non-stationary enough to resist agent optimization?
- Can AI systems improve themselves without external feedback?
- How do agents revise their own errors during autonomous architecture discovery?
- What makes self-modifying architectures learn their own update rules?
- How can agents evolve their own skills without human input?
- Does bounding textual edits prevent skill degradation better than free rewriting?
- What makes evolving the benchmark different from evolving the optimizer itself?
- What role does self-learning play in improving agent reasoning without annotation?
- Does self-play feedback improve skills created from the agent's own experience?
- Does learned evaluator co-evolution solve the problem of verifying hard-to-benchmark tasks?
- How many acceptable rewrites can recursive self-improvement sustain before returns diminish?
- Does swapping formal proofs for benchmarks change self-improvement safety?
- Can autonomous research agents outperform hand-tuned hyperparameter search?
- What distinguishes scaffold-level changes from parametric weight updates in self-improvement?
- Does removing static external utility break the formal guarantees of self-improvement loops?
- Can co-evolved critics truly circumvent static evaluator limitations in self-improvement?
- What capabilities can emerge from self-modification that the original agent lacked?
- Can AI systems generate and refine their own objective functions?
- How would a parametric self-improvement loop differ from a non-parametric one?
- How does compiling natural language goals into executable code enable objective evolution?
- How does controlling skill text edits prevent cascading failures in self-improvement?
- How would a bi-level agent restructure objective functions during discovery?
- Does AIDE2's single loop differ from bilevel autoresearch's nested loops?
- How do frozen executors with editable text state compare to end-to-end fine-tuning?
- Does AIDE2 archive rejected variants the way evolutionary approaches do for future reuse?
- Can moving or evolving objectives prevent misalignment in discovery agents?
- How does this scoped definition relate to the survey's open-ended recursive self-improvement?
- How does controlled utility evolution prevent the evaluator from becoming a new bottleneck?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- What makes an agent in an economic simulation self-evolving?
- Can a proposer agent actively surface a solver's weaknesses to prevent plateau?
- Does the preserve-and-extend contract alone drive the 17-point improvement?