Line of inquiry
Inquiring lines›What drives capability improvement…›What drives capability improvement…›this line of inquiry
Can code harness improvements rival direct model scaling for capability?
A broader line of inquiry — a family of 34 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 34
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can mid-tier models benefit more from harness improvements than frontier models?
- Can weaker models benefit equally from harness updates as stronger ones?
- How much does harness design contribute to reported model capability scores?
- Does harness scaling represent a fundamentally different path than model scaling?
- How much does harness design separate from model intelligence affect benchmark scores?
- Can mid-tier models benefit more from self-generated harness updates than others?
- How do prompt optimization and code harnesses compare for capability transfer?
- Why do useful harness updates often disappear during model evolution?
- Can weaker models match stronger ones by reorganizing harness-side components?
- What happens when different harnesses project the same model?
- Do harness improvements eventually internalize into core model behavior over time?
- How does harness structure affect planner token efficiency compared to model size?
- How does editing the harness layer differ from updating model weights?
- Do models co-adapt their harnesses to specific executor strengths?
- What cognitive burdens should move from model parameters into harness infrastructure?
- What makes a harness low-friction for model strategy?
- What makes a harness a first-class object rather than invisible scaffolding?
- What causes weak models to fail at activating harness artifacts?
- Why does harness benefit capacity peak at mid-tier models, not frontier scale?
- Can runtime behavior mapping help localize harness deficiencies?
- How should harness scaffolding be treated as a first-class object?
- Which foundation model tiers most benefit from harness updates?
- What distinguishes the fast scaffold learning loop from parametric model weight updates?
- Does harness benefit depend on which model tier you use?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- What happens when you project the same model onto different harnesses?
- How much of AI improvement comes from tools versus model capability?
- How does AI system design amplify model capabilities beyond the weights?
- What makes skills worth externalizing into a persistent harness?
- Can scaffold-only modifications achieve lasting gains without updating the foundation model weights?
- What makes API-based scaffolding more trustworthy than direct model access in high-stakes domains?
- How much does executor choice change a harness's actual performance?
- What should an external contract for model improvement actually contain?
- How much external enforcement does each model need before utility drops?