INQUIRING LINE

Train an AI only on expert examples and it never sees its own mistakes, so can it outgrow the people who picked them?

Can agents learn beyond the boundaries of their training data curators?

This explores whether AI agents can become more capable than the people who assembled their training examples, and what has to change for that to happen.


This explores whether AI agents can outgrow the people who assembled their training examples. The corpus starts from a clear diagnosis: agents trained only on expert demonstrations never act in a real environment during training, so they never see their own mistakes. Their ceiling is set by what the curators thought to include, not by what the model could do Can agents learn beyond what their training data shows?. Almost every other note here offers a different way to break through that ceiling. The common thread is to replace curated examples with consequences the agent produces itself.

The simplest version is to let the agent learn from what its own actions cause. One line of work treats the next state of the environment as the supervision signal, with no reward and no expert. That approach matches expert-dependent methods with half the data, and it gives reinforcement learning a better starting point Can agents learn from their own actions without external rewards?. Other approaches skip retraining entirely. Reflexion has agents write short verbal diagnoses of their own failures and read them back in later attempts. A plain pass/fail signal works best because it leaves no room to rationalize a failure Can agents learn from failure without updating their weights?. AgentFly extends this into a full memory system that improves agent behavior without touching the model's weights Can agents learn continuously from experience without updating weights?. VOYAGER stores working skills as code and combines them into new ones, while an automatic curriculum keeps pushing it into territory nobody planned Can agents learn new skills without forgetting old ones?. There is even evidence that learning can happen in the world itself: RL agents end up using their surroundings as a kind of memory without being told to Do RL agents accidentally use environments as memory?.

The more surprising move is to replace the human curator with a trained one. SkillOS trains a separate curator model to edit an agent's skill library. The library gradually shifts away from long, generic entries toward strategies that work across many tasks, and the trained curator still helps when paired with different agents Can a separate trained curator improve skill libraries better than frozen agents?. FlowReasoner goes further and designs a new multi-agent setup for each individual query instead of reusing a fixed template Can AI systems design unique multi-agent workflows per individual query?. A survey on co-evolution describes this as a progression. First the agent's peers adapt, then the environment and feedback adapt, and finally the process of improvement itself evolves. It also warns that a single agent improving itself in an unchanging setting tends to stall Can agents evolve beyond the constraints humans engineer?.

The catch is that some of these escape routes bring the boundary back in a new place. Language 'world models' trained on more than 10 million agent trajectories can simulate environments well enough to beat training in the real ones on three benchmarks Can language models learn to simulate agent environments?. Likewise, LLMs can stand in for search engines during training Can LLMs replace search engines during agent training?. But a simulated world only contains what the simulating model already knows. The limit moves from the human curators' imagination to the model's own. So the short answer is yes, agents can learn past their curators, but only to the extent they face real feedback: actual environments, binary outcomes, other adapting agents. When the agent trains inside its own simulation, the ceiling rises but does not disappear.


Sources 11 notes

Can agents learn beyond what their training data shows?

Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.

Can agents learn from their own actions without external rewards?

Research across eight environments shows that agents can use future states from their own actions as supervision without external rewards, matching expert-dependent baselines with half the data and providing superior warm-starts for subsequent RL training.

Can agents learn from failure without updating their weights?

Reflexion demonstrates that unambiguous environmental feedback (success/failure) enables agents to write useful self-diagnoses and improve across episodes without parameter updates. The binary signal prevents rationalization, and keeping reflections uncompressed preserves their usability.

Can agents learn continuously from experience without updating weights?

AgentFly formalizes agent learning as a Memory-augmented MDP with three memory modules (case, subtask, tool) that enable credit assignment and policy improvement entirely through memory operations. The approach achieved 87.88% on GAIA validation without modifying LLM parameters.

Can agents learn new skills without forgetting old ones?

VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.

Show all 11 sources
Do RL agents accidentally use environments as memory?

Mathematical proof shows that environmental artifacts reduce information needed to represent history in RL agents. Path-following agents naturally develop memory-like behavior through standard reward optimization, satisfying situated cognition criteria without explicit memory objectives.

Can a separate trained curator improve skill libraries better than frozen agents?

SkillOS shows that separating a trainable curator from a frozen executor, grouped by task streams, causes skill repositories to shift from generic verbose additions toward actionable execution logic and cross-task meta-strategies. The trained curator generalizes across different executor backbones and domains.

Can AI systems design unique multi-agent workflows per individual query?

FlowReasoner demonstrates that meta-agents trained with reinforcement learning and external execution feedback can generate unique multi-agent architectures for each user query, optimizing across performance, complexity, and efficiency—moving beyond fixed task-level workflow templates.

Can agents evolve beyond the constraints humans engineer?

A survey framework organizes co-evolving systems into three stages that progressively remove human engineering: dynamic peers first, then adaptive environments and feedback, finally the evolution mechanism itself. Single-entity self-improvement stalls in static contexts; co-evolution supplies adaptive pressure across multiple components.

Can language models learn to simulate agent environments?

Qwen-AgentWorld demonstrates that native language world models trained via next-state prediction on 10M+ trajectories outperform real-environment training on three benchmarks and transfer across seven domains, positioning next-state prediction as a foundation objective for agents.

Can LLMs replace search engines during agent training?

ZeroSearch and SSRL demonstrate that LLMs can generate relevant documents and search results from internal knowledge, with 14B simulators matching or exceeding real search engines. Curriculum degradation and test-time scaling optimize this approach for training without API costs.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.