Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›What factors determine agentic sys…›this line of inquiry
Can harness architecture and protocols provide agent reliability without model scaling?
A broader line of inquiry — a family of 49 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 49
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can smaller models trained for execution handle the failure modes that stop current agents?
- How does agent reliability emerge from memory and protocols instead of model scale?
- Where does agent reliability come from if not better tools?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- Can stopping rules extracted from past failures improve agent reliability without retraining?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
- Which harness dimensions most directly predict agent system reliability?
- What causes multi-turn agent failures: weak memory control or missing knowledge?
- What does error recovery look like across different agent architectures?
- Why do completion-mode strengths not transfer to agentic settings?
- Why does externalized state beat parameter scaling for agent reliability?
- Can slower development eliminate the risk of failure in agentic systems?
- How does the agentic layer amplify individual agent failure modes?
- What would an architecture that makes violations unavailable rather than unchosen look like?
- What architectural changes make violations unavailable rather than merely discouraged?
- What makes some model capabilities reliable while others remain brittle?
- Does adding capability without improving detection reduce overall system reliability?
- Can semantic audit layers attribute failure mechanisms to infrastructure-level state changes?
- Why does increased model capability make detection harder in delegated workflows?
- What distinguishes domain-specific failure modes from general model limitations?
- What makes action-producing models fail in ways text models typically do not?
- What degradation patterns emerge as relay length increases in delegated tasks?
- How do semantic discovery and tool selection failures differ between MCP and A2A?
- How do agentic systems recover when specialized models operate outside their scope?
- Why do long-horizon agents fail when their models can solve individual steps?
- Which eleven failure modes emerge from agentic layers in realistic deployment?
- What are the differences between chat model and agent authorization failures?
- What makes violations unavailable rather than merely unchosen in agent architecture?
- Why do frontier models corrupt more documents than weaker models during workflows?
- How does bounded committed state prevent multi-turn agent failures better than transcript replay?
- Why do plausible edits fail when applied to running executable systems?
- How should versioning and rollback govern the fast scaffold update loop?
- Why do frontier models corrupt documents while weaker models delete them?
- How does workflow scale change the failure modes of frontier models?
- Do architectural changes or training fixes better prevent agreement failures?
- Why does MCP's portability come with determinism failures in production workflows?
- Which agent architectures consistently outperform base models on hard prediction questions?
- Which failure modes dominate when models handle underspecified requests?
- What are the fourteen failure modes in deep research agents?
- Why do weak belief tracking and conservative actions trap agents in low-information states?
- Why is complex UI navigation the hardest agent failure mode?
- Can end-to-end models maintain debuggability without modular components?
- How does model tier affect whether errors delete or corrupt document content?
- What mechanism explains why context management prevents overflow failures most?
- Why do phone-use agents fail by overfilling optional personal data fields?
- Where does an agent's risk come from across its components and sequence?
- Can agents escape weak belief tracking and conservative action selection traps?
- What four domain properties make self-healing failure loops actually work?