INQUIRING LINE

If you tweak an AI's prompts or tiny internal signals instead of retraining it, does its stated reasoning still tell you what it's really doing?

Do weight-free agent edits keep chain-of-thought observations meaningful?

This explores whether improving an AI agent without retraining it (by changing its harness, prompts, or internal activations instead of its weights) leaves its visible reasoning trace an honest view of what it is actually doing.


This explores whether improving an agent without retraining it (by changing the scaffolding around the model, steering its prompts, or editing its hidden activations) keeps its chain-of-thought trustworthy as a view into how it reaches answers. The collection has no study that tests this head-on. What it does have is a set of findings that, read together, suggest the question is less about weights and more about whether the reasoning trace was telling us much to begin with.

The strongest reason to hope that weight-free edits help is that weight changes can do visible damage. Fine-tuned models produce reasoning chains that matter less to their final answers. You can cut the chain short, paraphrase it, or swap in filler text, and the answer more often stays the same, which suggests the reasoning has become a performance rather than the work itself Does fine-tuning disconnect reasoning steps from final answers?. On the other hand, harness scaling lifts frozen models on Terminal-Bench purely by rebuilding the execution system around them, and the same runbook carries over to newer models unchanged Can execution harnesses lift model performance without retuning weights?. If the weights don't move, the model's relationship between its reasoning and its answers shouldn't move either. Whatever faithfulness the base model had, it keeps.

The catch is that the base model's faithfulness was shaky before anyone touched it. Chain-of-thought examples that are logically invalid work almost as well as valid ones, which means models pick up the form of reasoning more than real inference Does logical validity actually drive chain-of-thought gains?. Trace length tracks how close a problem is to the training data, not how hard it is Does longer reasoning actually mean harder problems?. Chain of Draft matches full reasoning with 7.6% of the tokens, so most of a typical trace is style and documentation, not computation Can minimal reasoning chains match full explanations?. So keeping the weights frozen preserves a trace that was already partly decorative.

"Weight-free" also isn't one thing. Some weight-free methods edit the reasoning directly. One test-time intervention method found that verification and backtracking steps get very little attention from the model's later steps, and it removes about 75% of reasoning steps without losing accuracy Can reasoning steps be dynamically pruned without losing accuracy?. That is a useful lesson for anyone reading traces: the steps that look most like careful thinking, such as "let me double-check," may be the ones the model barely uses. Representation finetuning (ReFT) goes further. It leaves the weights frozen but rewrites hidden activations mid-computation Can editing hidden representations beat weight updates for finetuning?. That changes what the model computes as directly as fine-tuning does, so there's no reason to assume it protects faithfulness. Nobody in the collection has checked.

The practical point: if you want to know what an agent did, its narrated reasoning may be the wrong thing to watch. Agent-based evaluation that gathers outside evidence, such as files, outputs, and actual behavior, cut judge disagreement roughly 100-fold compared with an LLM judging text alone Can agents evaluate AI outputs more reliably than language models?. A harness-only edit makes that kind of checking easier, because the harness is where actions are logged. The more reliable signal may be what the agent does, not what it says it's thinking.


Sources 8 notes

Does fine-tuning disconnect reasoning steps from final answers?

Three faithfulness tests show fine-tuned models generate reasoning chains that less reliably influence final outputs. Early termination, paraphrasing, and filler substitution all produce invariant answers more often after fine-tuning, suggesting reasoning becomes performative rather than functional.

Can execution harnesses lift model performance without retuning weights?

StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.

Does logical validity actually drive chain-of-thought gains?

Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.

Does longer reasoning actually mean harder problems?

Controlled A* maze experiments show trace length correlates with difficulty only in-distribution but decouples entirely out-of-distribution. Trace length primarily reflects recall of training schemas, not adaptive computation.

Can minimal reasoning chains match full explanations?

Chain of Draft achieves equivalent accuracy to standard chain-of-thought on arithmetic, symbolic, and commonsense tasks while using only 7.6% of tokens. The 92.4% of removed tokens served style and documentation, not computation.

Show all 8 sources
Can reasoning steps be dynamically pruned without losing accuracy?

The PI framework categorizes reasoning into six types and uses attention maps to identify that verification and backtracking steps receive minimal downstream attention. Selecting only high-attention steps preserves accuracy while cutting reasoning length substantially.

Can editing hidden representations beat weight updates for finetuning?

ReFT learns task-specific interventions on frozen model representations rather than updating weights, with LoReFT (low-rank linear subspace variant) dramatically outperforming LoRA across reasoning, instruction-following, and NLU benchmarks while using far fewer parameters.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.