Will the tricks an AI agent learns while coding also make it better at tasks that have nothing to do with writing code?
Do agent improvements discovered on code tasks transfer to non-coding domains as well?
This explores whether the tools, workflows, and scaffolding that agents discover or refine while working on programming tasks keep helping when the agent moves to work that isn't about writing software, as opposed to only helping on more coding.
This explores whether improvements agents pick up on coding tasks carry over to non-coding work, or whether they only help with more code. The short answer from the corpus: there is strong evidence that these improvements travel across models and programming languages, some evidence that they reach neighboring technical fields, and little direct evidence that they reach clearly non-coding domains. The most-cited transfer result is Sakana AI's Darwin Gödel Machine. Improvements it found to its own tools and workflows while training on Python carried over to Rust, C++, and Go, and to different foundation models Do agent improvements discovered in one model transfer to others?. That is real portability, but every one of those destinations is still a coding task.
The furthest reach in the corpus is AIDE2. Its gains held up on four benchmarks it had never been tuned on, including physics-based weather forecasting, which lay outside the tasks used to select its improvements Do AIDE2's improvements transfer to unseen tasks?. Weather forecasting is a scientific domain, but the agent still does the work by writing and running code. Harness work shows a similar pattern. A fixed runbook built around frozen models lifted scores on Terminal-Bench across several models without any retuning Can execution harnesses lift model performance without retuning weights?. That shows the improvement isn't tied to one model. It doesn't show it works outside the terminal.
The corpus suggests a different way to see the question. Code may not be a domain at all. It may be the medium an agent thinks through. One note argues that code is uniquely executable, inspectable, and stateful, which lets it hold an agent's reasoning, its actions, and its checks in one loop Can code serve as the operational substrate for agent reasoning?. If that's right, a coding-born improvement transfers whenever a non-coding task can be turned into code, and the AIDE2 weather result looks like that kind of reframing. The improvements most likely to travel are the ones that live outside the model's weights: memory, skills, and protocols that take repeated work off the model Where does agent reliability actually come from?. Researchers favor these fast, cheap scaffold updates over retraining partly because they are easy to reverse and reuse Do self-improving agents really split into two distinct loops?.
The best evidence for crossing domains comes from methods that never started in code. Skills mostly work as procedural anchors that steady an agent's actions, not as packets of missing facts Do skills teach procedures or inject missing facts?. Procedures are the kind of thing that can carry over. A trained skill curator in SkillOS learned cross-task strategies that worked across different executor models and domains Can a separate trained curator improve skill libraries better than frozen agents?. Reusable sub-task routines gave large gains on web navigation, and the gains grew as the training and test tasks diverged Can agents learn reusable sub-task routines from past experience?. A language world model trained to predict the next state transferred across seven domains Can language models learn to simulate agent environments?.
The gap the corpus leaves open is the experiment that would settle this: take an improvement discovered on code and test it on something like customer support, writing, or negotiation. Nothing in this set does that. It also hints at why transfer might be limited. Agents learning only from expert demonstrations are capped by what the curators imagined Can agents learn beyond what their training data shows?. Group-evolving agents improved by pooling code patches and execution traces Does sharing experience across agents beat isolated evolution?, a signal that coding provides cheaply and most other domains don't. Coding may be where agent improvements are easiest to discover because it gives automatic, checkable feedback. Whether those improvements generalize beyond code may depend less on the improvement itself and more on whether the new domain gives back that kind of feedback.
Sources 12 notes
Sakana AI's Darwin Gödel Machine discovered improvements to agent tools and workflows that transferred to different foundation models (Claude 3.5 Sonnet, o3-mini, Claude 3.7 Sonnet) and to programming languages outside its training domain (Python-trained agents improved on Rust, C++, Go), suggesting the improvements target portable agent design rather than model-specific exploits.
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
Research shows code uniquely enables agent reasoning, action, and verification by being simultaneously executable, inspectable, and stateful. This unified code-centered loop improves reasoning and verification together compared to natural-language or prose-based approaches.
Research shows reliable LLM agents externalize three cognitive burdens—memory (state persistence), skills (procedural components), and protocols (structured interaction)—into a harness layer rather than relying on model scale alone. The harness unifies these externalities and eliminates the need for the model to solve the same problems repeatedly.
Show all 12 sources
A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.
Analysis of 8,135 trials shows procedural anchoring accounts for 65.7% of skill cases versus 4.5% for knowledge injection. Skills fail when retrieved incorrectly, invoked out of context, or followed too rigidly.
SkillOS shows that separating a trainable curator from a frozen executor, grouped by task streams, causes skill repositories to shift from generic verbose additions toward actionable execution logic and cross-task meta-strategies. The trained curator generalizes across different executor backbones and domains.
Agent Workflow Memory induces sub-task routines at finer granularity than full tasks, abstracts example-specific values, and compounds them hierarchically. This produces 24.6% relative gain on Mind2Web and 51.1% on WebArena, with larger gains as train-test gaps widen.
Qwen-AgentWorld demonstrates that native language world models trained via next-state prediction on 10M+ trajectories outperform real-environment training on three benchmarks and transfer across seven domains, positioning next-state prediction as a foundation objective for agents.
Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.
Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Demystifying Agent Skills: Why They Work-Until They Don't
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Self-Improvements in Modern Agentic Systems: A Survey
- Hyperagents
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents