Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›What factors determine agentic sys…›this line of inquiry
How do agent-learned skills transfer and improve across different tasks?
A broader line of inquiry — a family of 72 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 72
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do learned workflows transfer between different agents with minimal accuracy loss?
- Can individual skills improve through reuse and accumulate experience across tasks?
- Can skill repositories evolve toward execution-oriented refinement over time?
- Can agent skills move from prompts to trainable parameters?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- Can agent-authored skill libraries compound autonomy gains over time?
- Do weight-space skills lose detail compared to textual skill descriptions?
- How does real tool integration change what agents learn compared to simulated tools?
- What role does environment diversity play in preventing agents from overfitting to curator imagination?
- Can agents improve from deployment signals without explicit human annotation?
- Can context management policies transfer across agents of similar capability levels?
- Why do trajectory-based skills fail to transfer across different environments and use cases?
- What infrastructure decouples generation from training in asynchronous agent loops?
- Can skills learned from interaction trajectories outperform skills from static repositories?
- How can agent data flywheels improve task quality iteratively?
- Does a tight learning-rate bound prevent skills from escaping poor starting points?
- What domain properties determine whether causal rules transfer to new agents?
- Why do generic skill descriptions evolve into execution-oriented ones?
- Can RL-trained policies outperform text-space optimizers for evolving skill repositories?
- Can agents escape training data distributions without expensive real-world interaction?
- What training method supports dynamic tool discovery in long-horizon agents?
- How do human-agent systems incorporate diverse feedback into model behavior?
- What stops evolved agent behaviors from generalizing beyond specific tasks?
- Can next-state supervision work across different agent interaction types like conversations and tool calls?
- Can RL-trained meta-agents match or exceed manually designed workflows?
- Do task-level outcomes provide sufficient supervision for harness evolution?
- Can agents teach each other skills without human supervision?
- How does cross-agent supervision expand the set of convergent initial conditions?
- Can small numbers of curated demonstrations produce emergent agentic behavior?
- How do agents automatically generate suitable learning tasks based on current capability?
- How do complexity, diversity, and real-world fidelity interact in agent training?
- Can applicability conditions be preserved automatically when agents reflect on trials?
- Can agents learn to use scaffolding structure the way they learn token weights?
- Can tool adaptation work without freezing the agent in the loop?
- Can curriculum approaches teach agents when to stop exploring?
- Should optimal context budgets scale with agent competence or task complexity?
- What makes an agent mechanism reusable versus benchmark-specific?
- How does effective feedback retention govern long-horizon agent reliability?
- How do task stream groupings provide long-horizon learning signals for curation decisions?
- Does intentionally varying environment properties isolate causal effects on agent performance?
- How much does external context management transfer across similar capability agents?
- How should GUI agents remember patterns across different software environments?
- How does next-turn reward optimization contribute to agent passivity?
- Can combinational creativity alone drive open-ended learning in agents?
- How do fast and slow timescales enable continual agent adaptation?
- Do dynamic environments enable different kinds of agent-environment coevolution?
- What role does online RL play in scaling GUI agents?
- What prevents individual session learning from scaling system-wide?
- Does the 78-demonstration principle apply to other AI capabilities beyond agency?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- How does co-player diversity force agents to develop general adaptation?
- How does SDPO relate to agents learning from verbal reflection without parameter updates?
- Should role-play evaluation measure agent ability or user-agent pair fit?
- How much does agent performance depend on demonstration quantity versus curation quality?
- How do agents decide which skills to chain together for a single task?
- Can agentic AI tools deliver productivity gains on learning tasks differently?
- What training difficulty and curriculum settings prevent instability in empathetic agent RL?
- What role does sequence model in-context learning play in multi-agent cooperation?
- Does situational awareness training increase agents' ability to detect real deployment?
- Should user simulators be trained via RL like agents or decomposed into trackable state components?
- Why do environments authored once encode only what their builders imagined?
- What specific qualities make some demonstrations more effective for agency training?
- Why should environment properties scale alongside agent complexity and real-world fidelity?
- How do you prevent stale reward signals when skills evolve during deployment?
- What properties of agent systems only become visible across multiple sessions?
- Can agents revise their beliefs predictably when presented with interventions?
- Why does delegation training help models that work alone?
- What makes behavioral cloning produce more persuadable but less aligned agents?
- Can context management be optimized for an agent without retraining or changing the model?
- How do parametric and non-parametric updates differ in agents?
- How do agent capabilities change across 25 relay rounds of interaction?
- How do diagnose-and-reshape loops compare to building new environments from scratch?