Line of inquiry
Inquiring lines›How do agents behave and coordinat…›How do agents in production pipeli…›this line of inquiry
How do persistent skill repositories improve agent reliability over time?
A broader line of inquiry — a family of 62 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 62
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can skill repositories evolve toward execution-oriented refinement over time?
- Are durable shared code artifacts better than per-task harness patches?
- How do agents decide which created code deserves long-term persistence?
- Where does agent reliability come from if not better tools?
- How do agents decide which created code should persist versus disappear?
- Do learned workflows transfer between different agents with minimal accuracy loss?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- Should production agents execute one tool or multiple tools per invocation?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- Why do production agents depend more on their surrounding pipeline than the model?
- How does agent reliability emerge from memory and protocols instead of model scale?
- What separates good workflow design from poor workflow design?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
- Which harness dimensions most directly predict agent system reliability?
- Why do a-priori procedural specifications fail as environments change and interfaces evolve?
- How do agents discover and select which tools to invoke?
- Does agent-side context control outperform external management on any task class?
- Can skill libraries prevent redundant narrow artifacts from proliferating?
- When should agent-created code be promoted into permanent harness infrastructure?
- Can deterministic function calls prevent agent failures better than protocol-mediated tool access?
- How should agents decide which created code is worth persisting?
- Can slower development eliminate the risk of failure in agentic systems?
- What permission models govern code execution within agent skills?
- How does the execution layer constrain agent performance in tool use?
- How much does external context management transfer across similar capability agents?
- Why do workflow abstractions fail in embodied agent environments?
- What lifecycle management prevents in-loop skill creation from bloating an agent?
- Can tool adaptation work without freezing the agent in the loop?
- Why do 85 percent of production agents avoid third-party frameworks?
- Can one-off agent code be safely promoted to durable infrastructure?
- How should AI skills be created and managed like software artifacts?
- Can skill validation through testing prevent unreliable programs from accumulating?
- How do specialized agent roles improve consistency in long-form writing?
- Why do rigid orchestration frameworks fail where generative environment specifications succeed?
- Which ecosystem conditions matter most for agent deployment success?
- How do agents decide which skills to chain together for a single task?
- Does adding capability without improving detection reduce overall system reliability?
- What makes persistent, shared code artifacts from agents hard to manage at scale?
- Should agent capability be optimized separately from general capability?
- Can open agent workflows be modeled as finite event lifecycles?
- Why do production AI agents deliberately stay simple and avoid frameworks?
- Can agentic AI tools deliver productivity gains on learning tasks differently?
- What ecosystem conditions must exist for agents to function as economic participants?
- How much does agent performance depend on demonstration quantity versus curation quality?
- Why do persistent AI systems require fundamentally different design than ad-hoc supporters?
- Does AIDE2 archive rejected variants the way evolutionary approaches do for future reuse?
- Can disposable agent-authored code be distinguished from reusable infrastructure?
- How should versioning and rollback govern the fast scaffold update loop?
- How do agentic systems recover when specialized models operate outside their scope?
- What properties of agent systems only become visible across multiple sessions?
- How should skills be trusted and installed on sharing platforms?
- Why should environment properties scale alongside agent complexity and real-world fidelity?
- Which agent architectures consistently outperform base models on hard prediction questions?
- What specific bookkeeping tasks can environments maintain more reliably than policies?
- How do skills authored in-loop validate faster than offline generated skills?
- What specific qualities make some demonstrations more effective for agency training?
- Why does MCP's portability come with determinism failures in production workflows?
- Why do high-level design guidelines fail to capture real-world deployment nuance?
- How do agents discover and construct new APIs from existing applications?
- Can curator modules trained on one executor transfer to entirely different agent backbones?
- Where does an agent's risk come from across its components and sequence?
- How do controllable simulators compare to population-level agent simulation approaches?