SYNTHESIS NOTE
Topics›Context Engineering›this note

Does task success reward alone teach summary self-consistency?

When training an agent to write working-state summaries, can outcome-based RL fix the problem where a summary's recorded state contradicts its proposed next action, without any explicit compaction-specific reward signal?

Synthesis note · 2026-10-08 · sourced from Context Engineering

AutoCompact trains a coding agent to decide when to invoke a compact() action, what working-state summary to write, and how to continue afterward, using judge-corrected trajectories for supervised fine-tuning (SFT) followed by outcome-based reinforcement learning (RL) with "binary task success as the only reward signal." On SWE-bench Verified and SWE-PolyBench Verified it raises pass rates by an absolute 9.2% and 5.0% over the base model. The paper's sharpest finding concerns what the two training stages teach differently: a summary can satisfy both "preserving the relevant state and specifying a next action" while still "proposing a next action incompatible with the state it records" — a failure the authors name "summary self-consistency," distinct from information coverage. Figure 5 pairs paraphrased summaries from AutoCompact-SFT and the RL-trained AutoCompact on the same task; only the RL version keeps its stated next action compatible with its own recorded state.

The authors attribute the gap to what each objective optimizes: "SFT imitates demonstrated summaries, which does not ensure that the recorded state constrains the next action, whereas RL rewards a summary only through the success of the actions that follow it." Because GRPO rollouts receive a single binary outcome reward tied to whether the final patch passes task tests, a summary that misleads the next action shows up only as a downstream failure — so the only signal available to improve self-consistency is task completion, not any explicit judgment of the summary's coherence. RL also sharpens the trigger and coverage behaviors SFT first establishes: compact() usage rises from 44.3% of tasks (SFT) to 58.5% (RL) while key-state omissions fall from 3.1% to 0.2% and next-action omissions fall from 8.2% to 2.2%, "without compaction-specific rewards."

This extends Can thinking traces be made reliably budget-controllable?, which found that raw compressed traces fail at budget control and need a reward-driven framework to become reliable — AutoCompact locates a comparable failure specifically in the cross-reference between a summary's recorded state and its proposed next step, and shows plain outcome reward, with no compaction-specific term, suffices to fix it. It contrasts with Can an external manager handle context for frozen agents?: AdaCoM trains a separate manager to compress a frozen agent's context, while AutoCompact folds the compaction decision into the same policy being trained on the task, so coding and compaction are optimized jointly rather than through a module boundary. It also corroborates Which coding harness components matter most in different conditions?: AutoCompact's own ablation, running the trained checkpoint with compact() calls skipped, shows the advantage of actually executing compaction is largest under tight budgets — 19.9% at $0.10 versus 1.9% at $4.00 — matching that note's finding that context management pays off most when windows are tight.

The excerpt reports self-consistency qualitatively through one paired example rather than a measured rate across tasks, so the size of the self-consistency gap between SFT and RL is not established numerically here, only its direction. RL training also used shorter 32K-token sequences than the 256K window tested at evaluation, which the authors flag as a resource constraint rather than a validated sufficient condition — so it remains a working assumption, not a demonstrated result, that self-consistency learned at short horizons transfers cleanly to much longer ones.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do autonomous agents misreport success on failed actions?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 152 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

AutoCompact's task-success RL reward teaches summary self-consistency that SFT imitation of compaction summaries does not