Line of inquiry
Inquiring lines›What enables authentic and grounde…›How should retrieval-augmented gen…›this line of inquiry
Does externalizing cognitive work and state improve agent reliability?
A broader line of inquiry — a family of 33 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 33
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does externalizing reasoning into harness artifacts improve agent reliability?
- Why does externalized state beat parameter scaling for agent reliability?
- How does external context control compare to agents managing their own state internally?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
- Can externalizing bookkeeping to a stateful harness replace internalized memory control?
- Where does agent reliability come from if not better tools?
- What persistent failures remain unsolved despite harness evolution efforts?
- Why do memory and feedback loops matter more than model size for agent reliability?
- What makes behavior localization the bottleneck in agent harness evolution?
- How should the surrounding agent system be designed to ground actions in reality?
- Why does the harness layer accumulate distributed behaviors over time?
- Why do a-priori procedural specifications fail as environments change and interfaces evolve?
- What makes some model capabilities reliable while others remain brittle?
- Why does externalizing bookkeeping raise effective feedback compute?
- What governance and safety measurements matter for deployed agent environments?
- What specific bookkeeping tasks can environments maintain more reliably than policies?
- How do agentic systems recover when specialized models operate outside their scope?
- Can skill validation through testing prevent unreliable programs from accumulating?
- How should versioning and rollback govern the fast scaffold update loop?
- What makes skills worth externalizing into a persistent harness?
- Why do persistent, resynchronized artifacts compound harness capability gains?
- How can harnesses externalize bookkeeping so models focus on semantic judgment?
- What training difficulty and curriculum settings prevent instability in empathetic agent RL?
- Why do high-level design guidelines fail to capture real-world deployment nuance?
- What execution-layer design prevents agents from passively reacting to environments?
- What safety protections work when simulators have access to real APIs?
- How do capability tracks and behavior tracks stay separable during skill deployment?
- Why does MCP's portability come with determinism failures in production workflows?
- Can end-to-end models maintain debuggability without modular components?
- How do virtual model instances preserve identity through load-balancing and failover?
- Does inspectable skill artifacts guarantee the behavior matches the person it claims to ground?
- Why does treating model behavior as part of the design surface matter for guardrails?