Why do AI agents that learn by copying experts never get past the experts, since they never try things themselves?
Why do imitation learning agents stay locked within human demonstration patterns?
This explores why agents trained to copy expert examples rarely get better than those examples, and what the corpus says about ways past that ceiling.
This explores why agents trained to copy expert examples rarely get better than those examples, and what the corpus says about ways past that ceiling. The short answer is that imitation learning never lets the agent touch the world while it is learning. It sees what experts did, but it never sees what happens when it tries something itself, including when it fails. So the agent's competence is capped by what the people who built the dataset thought to include, not by what the agent could learn (Can agents learn beyond what their training data shows?). If no one demonstrated a recovery from a particular mistake, the agent has no way to learn one.
There is a second limit. Copying tends to pick up the surface of the examples more than the reasoning behind them. One related finding comes from language models: models instruction-tuned on meaningless or even wrong instructions scored about the same as models tuned on correct ones. What carried over was a sense of what outputs should look like, not an understanding of the task (Does instruction tuning teach task understanding or output format?). That study is about instruction tuning, not agents, but it points the same way. Learning from demonstrations can teach the form of good behavior without teaching why it works.
The ways out all share one move: give the agent its own feedback. One approach treats the results of the agent's own actions as the training signal. It needs no reward and no expert, yet it matched expert-dependent baselines with half the data (Can agents learn from their own actions without external rewards?). Another has agents write short notes on their own failures and reread them on later attempts, without retraining (Can agents learn from failure without updating their weights?). A third has agents build a growing library of skills tested against the environment (Can agents learn new skills without forgetting old ones?). Self-play systems go further and create their own curriculum and judge (Can language models learn skills without human supervision?). A survey frames this as gradually removing human-built constraints, ending with the improvement process itself (Can agents evolve beyond the constraints humans engineer?).
The twist: human demonstrations don't become useless. Their job changes. In one driving study, thirty minutes of human data, used only as a light check on self-play, was enough to keep the learned policy compatible with human drivers. That is about 2,500 times less data than pure imitation needs (Can human data steer self-play RL toward human-compatible behavior?). In reasoning, imitating first and then refining with reinforcement learning beats either method alone. The imitation stage produces attempts reasonable enough for rewards to say anything useful about them (Does sequencing imitation then exploration training improve reasoning?).
The most surprising finding cuts the other way. In search agents, reinforcement learning, the usual cure for imitation's ceiling, narrows behavior: policies settle on a few strategies that maximize reward. Fine-tuning on diverse demonstrations keeps exploration broader (Does reinforcement learning squeeze exploration diversity in search agents?). So being locked into demonstrations and being locked into a reward are two versions of the same problem. The corpus suggests the best results come from combining the two, so that each covers the other's blind spot.
Sources 10 notes
Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.
Models trained on semantically empty or deliberately incorrect instructions achieve comparable performance to those trained on full correct instructions, achieving 43% vs random baseline 42.6%. The semantic content of instructions appears largely irrelevant; what transfers is knowledge of the output space.
Research across eight environments shows that agents can use future states from their own actions as supervision without external rewards, matching expert-dependent baselines with half the data and providing superior warm-starts for subsequent RL training.
Reflexion demonstrates that unambiguous environmental feedback (success/failure) enables agents to write useful self-diagnoses and improve across episodes without parameter updates. The binary signal prevents rationalization, and keeping reflections uncompressed preserves their usability.
VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.
Show all 10 sources
Ctx2Skill's three-role self-play loop manufactures missing feedback through internal signals: the Challenger escalates difficulty as curriculum, the Judge gives binary verdicts as reward, and both sides evolve via natural-language skill edits. Success requires balancing adversarial pressure against a generalization safeguard to prevent collapse.
A survey framework organizes co-evolving systems into three stages that progressively remove human engineering: dynamic peers first, then adaptive environments and feedback, finally the evolution mechanism itself. Single-entity self-improvement stalls in static contexts; co-evolution supplies adaptive pressure across multiple components.
Human demonstrations work best as a regularization layer on self-play rewards, not as primary training signals. Just 2500x less data than imitation learning achieves human-compatible driving policies in 15 hours on consumer hardware.
Running Supervised RL first to establish reasoning foundations, then RLVR to refine against verifiable rewards, substantially outperforms both methods in isolation. The imitation phase makes outcome rewards informative by creating reasonable rollouts the RL phase can then sharpen.
RL training compresses behavioral diversity in search agents through the same entropy collapse mechanism documented in reasoning—policies converge on narrow reward-maximizing strategies. SFT on diverse demonstrations preserves exploration breadth, suggesting diversity-preservation techniques are essential for RL search scaling.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Agent Learning via Early Experience
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Training Language Models to Self-Correct via Reinforcement Learning
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- Demystifying Agent Skills: Why They Work-Until They Don't
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Self-Questioning Language Models
- Sharpening Tax in Post-Training