Agent Harness
A subject the collection covers, read through 43 synthesis notes.
Can a routing harness generate its own training data automatically?
Explores whether the logs and signals produced by an agent routing system—which directs requests to appropriate model tiers—naturally contain the evidence needed to improve the models themselves through fine-tuning and distillation.
Does training editors on real outcomes beat prompting larger models?
Can a small model trained on whether its patches actually work outperform larger frontier models prompted to make the same edits? This matters because it tests whether feedback beats raw capacity for runtime system modification.
Does self-editing through reviewed commits improve agent performance?
Ouroboros evolves its own prompts, tools, and core code through a reviewed commit process and reports top benchmark scores. But without comparing to a frozen version of itself, the contribution of self-evolution versus initial design or model capacity remains unclear.
Can orchestration layers make coding agents more auditable?
Does wrapping a fixed coding agent in state tracking and skill libraries improve research auditability and completeness without replacing the agent itself?
Can external state caches let models solve harder problems?
Explores whether organizing model state across weights, context, persistent memory, and disk—rather than relying only on weights and tokens—expands what models can do. Matters because long-horizon tasks may need more state than a model can hold internally.
Should safety harnesses be customized for each deployment?
Can a single safety harness design work across different models and domains, or does each deployment need its own tuned version? Understanding this matters for scaling AI safety practices efficiently.
What are the three distinct layers of agent code?
Does separating agent code into model capabilities, system harness, and agent-created artifacts help explain why agentic systems fail and where to intervene for improvement?
How should we measure agent system performance beyond task success?
Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?
Where does agent reliability actually come from?
Exploring whether LLM agent performance depends on larger models or on thoughtful system design choices like memory, skills, and protocols that shift cognitive work outside the model.
What happens to code that agents create and then share?
Agent-authored code artifacts that persist across tasks and multiple agents remain poorly understood. The open questions cluster around what should be retained versus discarded, and how shared state stays consistent when multiple agents collaborate.
Does arbitrary code execution alone capture exploit progress?
ExploitGym scores only working code execution, but exploitation involves reaching intermediate primitives like memory read/write first. Does this top-step-only metric miss meaningful partial progress that defenders should care about?
Does harness self-improvement memorize tasks instead of learning broadly?
When agents automatically edit their own prompts and tools based on task feedback, do those improvements generalize to new domains or just fit the training tasks? This matters because overfitting at the harness level could hide real capability gains.
Why is finding distributed behavior code so hard?
When developers need to modify agent harnesses, they struggle to locate all the code implementing a target behavior because behaviors are scattered across files and stages while requests describe what to do, not where to look.
Can code serve as the operational substrate for agent reasoning?
Explores whether code functions not just as LLM output but as the executable medium through which agents reason, act, and verify progress. This reframing treats code as infrastructure rather than deliverable.
Which coding harness components matter most in different conditions?
Can individual harness components—planning, context management, action space—be evaluated separately rather than as a package? This matters because practitioners need to know which components to prioritize given their constraints.
Can context quality alone predict how agents will behave?
Can we score the quality of an agent's context independently—its instructions, tools, knowledge, guardrails—and use that score to forecast whether the agent will fail or succeed, without observing its actual behavior?
Do harness edits learn reusable strategies or memorize task fixes?
When meta-agents evolve harnesses iteratively, do the persisted edits encode transferable procedures that solve new problems, or do they mostly cache shortcuts for already-solvable tasks? This matters because it determines whether harness evolution genuinely expands capability.
Do cybersecurity benchmarks actually measure exploitation?
Frontier models score well on vulnerability finding, patching, and CTF challenges, but does that success tell us whether they can convert vulnerabilities into real attacks? The paper argues exploitation—turning a bug into actual impact—remains under-evaluated.
Does measuring exploit capability help or harm defense?
Exploitation benchmarks can support defenders and attackers equally. How should we evaluate capabilities with unavoidable dual-use potential, and what safeguards make evaluation itself defensible?
Why does exploitation test multiple reasoning demands at once?
Exploitation tasks layer memory reasoning, runtime adaptation, and long-horizon planning into a single challenge. Understanding how these demands interact helps diagnose which capability limits agent performance.
What causes failures in exploitation benchmarks?
Benchmark failures may come from safety refusals, tool misuse, or impossible tasks rather than lack of capability. This matters for assessing how dangerous AI agents could actually be.
Can language models build and maintain their own agent harnesses?
This explores whether an LLM's ability to create and revise its own execution infrastructure is a distinct skill from solving tasks within someone else's harness, and whether current evaluations overlook this capability.
How should we measure gains from automatic harness evolution?
Harness evolution itself runs a search loop, so reported improvements might come from more search rather than better design. What's the right way to measure whether the harness itself actually improved?
Can task state management alone improve long-horizon agent performance?
Does keeping verified task state outside the execution context, rather than embedded in a growing context window, help long-horizon agents succeed more often? This matters because it challenges assumptions about where bottlenecks actually occur.
Can harness modules improve separately from benchmark data?
Does evolving harness components independently on out-of-distribution data, using contrasted success and failure trajectories, help distinguish reusable improvements from task-specific overfitting? This matters because current methods conflate general gains with benchmark adaptation.
Can skills work better as weights than as prompts?
Most agent systems store skills as text in prompts, but this inflates token costs and degrades model performance. Could compiling skills into trainable weight-space adapters instead offer a better trade-off between efficiency and capability?
Can distilled skills close the gap in ML research agents?
ML research agents have strong models and planning harnesses, but lack domain-specific operational knowledge. Can compact, verified skills extracted from repositories and papers fill that gap and improve agent performance?
Can person-grounded skills remain auditable without hidden prompt state?
Explores whether treating extracted expertise as versioned files—rather than persona prompts—enables meaningful accountability over person-grounded knowledge. Matters because audit trails determine whether captured skills can be corrected, rolled back, or safely withheld.
How should agents route across thousands of skills?
As skill libraries grow, should routing focus on selecting one skill or composing many? This explores whether decomposition and chaining creates better task execution than single-skill selection.
Can explicit behavior maps help weaker planners compete with stronger models?
Explores whether organizing harness repositories around runtime behavior—rather than relying on model inference—can narrow the capability gap between weaker and stronger planning models, and whether this reduces computational overhead.
Can agent harnesses be automatically optimized across many environments?
Explores whether scaling auto-research loops across diverse harness environments can discover mechanisms that reduce token use without sacrificing task performance, and whether such discoveries generalize.
Can externalized bookkeeping let smaller search agents beat larger ones?
Does offloading routine record-keeping to an environment harness free RL policies to focus on semantic search decisions, and can this approach outperform larger searchers with fewer parameters?
Can frozen models improve by evolving their harnesses?
DarwinX reports 17-point gains from selecting harness variants while keeping model weights frozen. The question is whether this improvement comes from population-level selection, the non-regression contract, the archive mechanism, or some combination of the three.
Can source code replace experience as skill raw material?
Existing skill synthesis relies on agent trajectories or documents, each with limitations. Could static code repositories serve as a more reliable, scalable foundation for deriving reusable procedural skills?
Do skills teach procedures or inject missing facts?
This research explores whether skills help agents by providing procedural structure versus supplying new information. Understanding this distinction clarifies when and why skills improve performance.
What blocks skill retrieval in task decomposition?
When routing queries across a skill library, does the granularity of task decomposition determine whether retrieval can succeed? This explores whether fixing decomposition precision unlocks better skill matching.
Can scarcity of solutions protect benchmarks from data contamination?
ExploitGym lacks ground-truth exploits for many tasks, which might prevent models from memorizing solutions during training. But does difficulty-based protection actually hold up, or does it degrade once solutions become public?
Do stronger models always evolve harnesses better?
We explore whether base model capability predicts both the ability to write useful harness updates and the ability to benefit from them. The answer reshapes how we should allocate capability in self-evolving agent systems.
What makes agent memory quality better than storage capacity?
If agents need better memory, should we focus on adding storage or improving what gets kept? This explores why curation and selective forgetting matter more than raw capacity for reliable agent performance.
Can conversation patterns predict coding outcomes better than prior skill?
Researchers tested whether usage patterns automatically derived from human-AI coding sessions could predict success beyond what prior achievement explains. This matters for developing teachable guidelines as AI capabilities keep changing.
Which security protections actually slow down agent exploits?
ExploitGym varies defenses across 898 real-world instances to isolate how each protection affects agent performance. Understanding which defenses matter most to agents versus humans is critical for defenders.
Can agents learn to work reliably through environment and coordination scaling?
Does training agents in diverse, verifiable environments and teaching them to coordinate tasks produce sustained capability on real-world work? This matters because general-purpose models reason well but struggle with stateful, recoverable execution across tools.
Can wrapping environments reshape how agents learn without breaking verifiers?
Does adding a programmable layer around static environments let them adapt to agent weaknesses while preserving their original correctness checks? This matters because hand-built environments quickly become limiting as agents improve.