INQUIRING LINE

Fix how an AI agent handles its terminal and merges multiple tries into one answer — does that fix still work if you swap in a different AI model?

Do shell-state handling and answer aggregation fixes transfer across different models unchanged?

This explores whether harness-level fixes, such as how an agent keeps track of its shell or terminal session and how it combines several candidate answers into one, carry over unchanged when you swap the underlying model, or whether each model needs its own tuning.


This explores whether fixes made to the scaffolding around a model (how it keeps track of its shell session, how it combines several attempts into a final answer) work unchanged when you plug in a different model. The short answer: the collection has no paper that tests this transfer directly. It does contain several findings that point in the same direction. Harness fixes tend to transfer best when they take unstable work away from the model, and worst when they depend on how a particular model behaves.

The strongest evidence comes from the idea of separating harness failures from model failures. One long-horizon agent design keeps state at four levels: the model's weights, its context window, a persistent coding session with helper agents, and history saved to disk. Part of the reason for this layering is so that when something breaks, you can tell whether the scaffold or the model was at fault Can external state caches let models solve harder problems?. That separation is a precondition for transfer: a shell-state fix that lives entirely in the harness should, in principle, carry over to any model. The same note warns, though, that individual components weren't tested in isolation, so the claim about where the gains come from hasn't been fully checked. A related result shows a stronger model building inference-time harnesses that nearly doubled a weaker model's scores. It worked mainly by moving unstable reasoning into deterministic code and routing each task to a fixed procedure, not by asking for more reasoning Can a stronger model lift a weaker one at test time without retraining?. The pattern is that the more of the fix lives in code rather than in the model's habits, the more portable it is likely to be.

The counter-pressure comes from research on prompt sensitivity. How much a model's output changes when you rephrase the input tracks how confident that model is. Larger, more confident models barely move, while less confident ones swing a lot Does model confidence predict robustness to prompt changes?. Separately, models respond to how common a phrasing was in their training data rather than to what it means, so two prompts that mean the same thing can give different results Why do semantically identical prompts produce different LLM outputs?. Many harness fixes are partly prompt fixes: instructions about when to check the working directory, or how to format an answer so an aggregator can parse it. These findings suggest such fixes will behave differently from model to model, and will be most fragile on smaller or less confident models.

Answer aggregation has its own wrinkle. Majority voting and similar schemes assume each attempt is a fair sample. In diffusion-based language models, answer confidence settles early while the reasoning keeps changing Can reasoning and answers be generated separately in language models?. This means the point at which an answer is safe to read off depends on the model's architecture, so an aggregation rule tuned for one kind of model may not suit another. Shell handling has a similar problem. An agent trained to search by issuing grep commands over raw text learned stable command usage only after supervised examples followed by reinforcement learning Can direct corpus search beat embedding-based retrieval?. That suggests reliable shell behavior is partly something a model learns, not just something the harness provides.

The practical takeaway is to sort fixes by where they live, not by what they target. Deterministic fixes, like saving state to disk, restoring the working directory in code, or parsing answers with a program, should transfer close to unchanged. Fixes that depend on the model reading an instruction a particular way should be retested for each model. If you want this tested directly, the collection doesn't have it yet. The right experiment would apply the same harness change to several different models and measure the gain on each.


Sources 6 notes

Can external state caches let models solve harder problems?

Prime Agent organizes persistent state in four levels (weights, context, persistent REPL plus subagents, disk-backed history) to let models read and write addressable state beyond their instruction stream. The approach isolates harness failures from model failures and reported gains on ARC-AGI-3, though specific components remain unablated.

Can a stronger model lift a weaker one at test time without retraining?

A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.

Does model confidence predict robustness to prompt changes?

ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.

Why do semantically identical prompts produce different LLM outputs?

Cao et al. and Adam's Law show that semantically identical prompts with different sentence-level frequencies produce systematically different output quality. Higher-frequency phrasings win because models register statistical mass from pre-training, not meaning.

Can reasoning and answers be generated separately in language models?

ICE shows that bidirectional attention in diffusion LLMs enables in-place prompting—embedding reasoning directly in masked positions refined alongside answers. Answer confidence converges early while reasoning continues refining, allowing early-exit mechanisms to cut compute by 50% while maintaining accuracy.

Show all 6 sources
Can direct corpus search beat embedding-based retrieval?

GrepSeek trains agents to retrieve via executable shell commands over raw text, achieving better multi-hop performance on entity-constrained queries than dense embeddings. The approach scaffolds unstable search mechanics with supervised trajectories, then refines task-oriented behavior through reinforcement learning.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.