Line of inquiry
Inquiring lines›How do language models construct a…›How are AI-generated and human-wri…›this line of inquiry
Do harness improvements transfer across model scales or memorize shortcuts?
A broader line of inquiry — a family of 19 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 19
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can mid-tier models benefit more from self-generated harness updates than others?
- Can harness updates benefit agents equally across all model sizes?
- What components of agent scaffolding most impact domain-specific output quality?
- What cognitive burdens should move from model parameters into harness infrastructure?
- Can harness evolution be redirected toward distilling transferable procedures instead?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- Why do evolved harness edits mostly memorize rather than generalize?
- What happens when different harnesses project the same model?
- What causes weak models to fail at activating harness artifacts?
- Can smaller models produce skill updates as useful as frontier model updates?
- Do gains from harness-based agents transfer across different search benchmarks?
- What makes harnesses more tangled than other types of agent code?
- Does harness benefit depend on which model tier you use?
- How should harness scaffolding be treated as a first-class object?
- What feedback signals matter most during harness evolution search?
- How should we allocate model budget between evolvers and harness users?
- What happens when you project the same model onto different harnesses?
- What makes API-based scaffolding more trustworthy than direct model access in high-stakes domains?
- Can per-user adapters remain consistent without drifting or leaking?