Line of inquiry
Inquiring lines›How do agents behave and coordinat…›How can we measure and understand…›this line of inquiry
Can agents develop persistent skills that compound over time?
A broader line of inquiry — a family of 69 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 69
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does reasoning ability help agents learn from feedback faster?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- Can individual skills improve through reuse and accumulate experience across tasks?
- Can agent skills move from prompts to trainable parameters?
- Can agent-authored skill libraries compound autonomy gains over time?
- How can agents distinguish over-generalized lessons from genuinely useful long-tail knowledge?
- Can agents improve from deployment signals without explicit human annotation?
- Why do AI agents struggle with novel experiments but excel at routine tasks?
- What role does self-learning play in improving agent reasoning without annotation?
- Does a tight learning-rate bound prevent skills from escaping poor starting points?
- Can architectural changes like decoupling intent understanding help overcome next-turn reward limitations?
- What role does environment diversity play in preventing agents from overfitting to curator imagination?
- What infrastructure decouples generation from training in asynchronous agent loops?
- How does real tool integration change what agents learn compared to simulated tools?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- Do weight-space skills lose detail compared to textual skill descriptions?
- Can stopping rules extracted from past failures improve agent reliability without retraining?
- Can small numbers of curated demonstrations produce emergent agentic behavior?
- Why does externalized state beat parameter scaling for agent reliability?
- Why do agents fail to internalize value from informative observations?
- Can early experience replace external rewards as a learning signal?
- Can agents teach each other skills without human supervision?
- What makes trajectory quality matter more than one-shot task success?
- Does social scaffolding outperform purely intrinsic motivation for agent exploration?
- Can agents learn to distinguish helpful from misleading interventions?
- How do human-agent systems incorporate diverse feedback into model behavior?
- Can applicability conditions be preserved automatically when agents reflect on trials?
- Can episodic memory alone enable learning without parameter updates?
- What domain properties determine whether causal rules transfer to new agents?
- Can curriculum approaches teach agents when to stop exploring?
- Can agents learn to use scaffolding structure the way they learn token weights?
- How do agents ground their judgments in evidence instead of pattern matching?
- Can RL-trained policies outperform text-space optimizers for evolving skill repositories?
- How can agents evolve their own skills without human input?
- What stops evolved agent behaviors from generalizing beyond specific tasks?
- Can next-state supervision work across different agent interaction types like conversations and tool calls?
- Does self-play feedback improve skills created from the agent's own experience?
- How do agents automatically generate suitable learning tasks based on current capability?
- What equilibrium-selection problem does human data solve in multi-agent learning?
- Can simulation fidelity limit what agents learn from trained world models?
- What training method supports dynamic tool discovery in long-horizon agents?
- Can RL-trained meta-agents match or exceed manually designed workflows?
- Can combinational creativity alone drive open-ended learning in agents?
- Can AI models retain knowledge across changing environments without catastrophic forgetting?
- How do task stream groupings provide long-horizon learning signals for curation decisions?
- What happens when agents interact with environments and learn from their own mistakes?
- Can zero-weight drift through external memory replace parameter plasticity entirely?
- Can influence estimation identify the most valuable trajectories in agentic training?
- Why does credit assignment through memory rewriting avoid expensive LLM parameter updates?
- How do agents differ in caution versus persistence across low-information scenarios?
- How does SDPO relate to agents learning from verbal reflection without parameter updates?
- How does credit assignment drive agents to write information into environments?
- Does intentionally varying environment properties isolate causal effects on agent performance?
- Does the 78-demonstration principle apply to other AI capabilities beyond agency?
- What mechanisms let later agents inherit information left by earlier ones?
- Do dynamic environments enable different kinds of agent-environment coevolution?
- How do evaluative versus directive signals differ in next-state training?
- Can graph topology represent successful trajectory clusters more effectively than skill libraries?
- What role does private information play in distinguishing realistic from unrealistic agents?
- How do fast and slow timescales enable continual agent adaptation?
- What role does sequence model in-context learning play in multi-agent cooperation?
- How does co-player diversity force agents to develop general adaptation?
- Does adding survey data to interviews improve agent accuracy further?
- How do intrinsic motivation principles explain why generating novel challenges improves learning?
- What can agents learn from the brain's complementary learning systems?
- Should user simulators be trained via RL like agents or decomposed into trackable state components?
- Can agents revise their beliefs predictably when presented with interventions?
- When does simulated search outperform real search for agent training?
- What cognitive capabilities do agents need to internalize social feedback?