Line of inquiry
Inquiring lines›How do training choices shape mode…›How do optimization strategies aff…›this line of inquiry
Do evolved harness improvements generalize as reusable strategies or memorize?
A broader line of inquiry — a family of 35 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 35
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do mid-tier models benefit most from memorized harness fixes?
- Do evolved harness edits capture reusable strategies or task-specific memorization?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- How do different harness designs produce different agent behaviors from the same model?
- Can harness updates benefit agents equally across all model sizes?
- Why do evolved harness edits mostly memorize rather than generalize?
- Why does the harness layer accumulate distributed behaviors over time?
- What components of agent scaffolding most impact domain-specific output quality?
- Can harness evolution be redirected toward distilling transferable procedures instead?
- What persistent failures remain unsolved despite harness evolution efforts?
- Can mid-tier models benefit more from self-generated harness updates than others?
- What makes behavior localization the bottleneck in agent harness evolution?
- What makes harnesses more tangled than other types of agent code?
- Can harness evolution be redirected from memorization toward strategy distillation?
- What cognitive burdens should move from model parameters into harness infrastructure?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- How much of harness-evolution gain comes from matched test-time search budgets?
- How do agent-created code artifacts become part of harness infrastructure?
- What feedback signals matter most during harness evolution search?
- What causes weak models to fail at activating harness artifacts?
- What happens when different harnesses project the same model?
- Do gains from harness-based agents transfer across different search benchmarks?
- What makes a harness a first-class object rather than invisible scaffolding?
- Why do persistent, resynchronized artifacts compound harness capability gains?
- How should harness scaffolding be treated as a first-class object?
- How does editing the harness layer differ from updating model weights?
- Why does harness benefit capacity peak at mid-tier models, not frontier scale?
- What makes agent-initiated artifacts the underexplored frontier in harness engineering?
- How can harnesses externalize bookkeeping so models focus on semantic judgment?
- How should the surrounding agent system be designed to ground actions in reality?
- Does harness benefit depend on which model tier you use?
- What makes durable code artifacts more valuable than per-task harness patches?
- How should we allocate model budget between evolvers and harness users?
- What happens when you project the same model onto different harnesses?
- What execution-layer design prevents agents from passively reacting to environments?