Does task success reward alone teach summary self-consistency?
When training an agent to write working-state summaries, can outcome-based RL fix the problem where a summary's recorded state contradicts its proposed next action, without any explicit compaction-specific reward signal?
AutoCompact trains a coding agent to decide when to invoke a compact() action, what working-state summary to write, and how to continue afterward, using judge-corrected trajectories for supervised fine-tuning (SFT) followed by outcome-based reinforcement learning (RL) with "binary task success as the only reward signal." On SWE-bench Verified and SWE-PolyBench Verified it raises pass rates by an absolute 9.2% and 5.0% over the base model. The paper's sharpest finding concerns what the two training stages teach differently: a summary can satisfy both "preserving the relevant state and specifying a next action" while still "proposing a next action incompatible with the state it records" — a failure the authors name "summary self-consistency," distinct from information coverage. Figure 5 pairs paraphrased summaries from AutoCompact-SFT and the RL-trained AutoCompact on the same task; only the RL version keeps its stated next action compatible with its own recorded state.
The authors attribute the gap to what each objective optimizes: "SFT imitates demonstrated summaries, which does not ensure that the recorded state constrains the next action, whereas RL rewards a summary only through the success of the actions that follow it." Because GRPO rollouts receive a single binary outcome reward tied to whether the final patch passes task tests, a summary that misleads the next action shows up only as a downstream failure — so the only signal available to improve self-consistency is task completion, not any explicit judgment of the summary's coherence. RL also sharpens the trigger and coverage behaviors SFT first establishes: compact() usage rises from 44.3% of tasks (SFT) to 58.5% (RL) while key-state omissions fall from 3.1% to 0.2% and next-action omissions fall from 8.2% to 2.2%, "without compaction-specific rewards."
This extends Can thinking traces be made reliably budget-controllable?, which found that raw compressed traces fail at budget control and need a reward-driven framework to become reliable — AutoCompact locates a comparable failure specifically in the cross-reference between a summary's recorded state and its proposed next step, and shows plain outcome reward, with no compaction-specific term, suffices to fix it. It contrasts with Can an external manager handle context for frozen agents?: AdaCoM trains a separate manager to compress a frozen agent's context, while AutoCompact folds the compaction decision into the same policy being trained on the task, so coding and compaction are optimized jointly rather than through a module boundary. It also corroborates Which coding harness components matter most in different conditions?: AutoCompact's own ablation, running the trained checkpoint with compact() calls skipped, shows the advantage of actually executing compaction is largest under tight budgets — 19.9% at $0.10 versus 1.9% at $4.00 — matching that note's finding that context management pays off most when windows are tight.
The excerpt reports self-consistency qualitatively through one paired example rather than a measured rate across tasks, so the size of the self-consistency gap between SFT and RL is not established numerically here, only its direction. RL training also used shorter 32K-token sequences than the 256K window tested at evaluation, which the authors flag as a resource constraint rather than a validated sufficient condition — so it remains a working assumption, not a demonstrated result, that self-consistency learned at short horizons transfers cleanly to much longer ones.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do autonomous agents misreport success on failed actions?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can thinking traces be made reliably budget-controllable?
Raw thinking traces compress well but ignore budget targets and take shortcuts. Can reward optimization make them controllable and useful for deployment?
both show raw or imitated compression needs a reward-driven training signal to become reliable, not merely compact
-
Can an external manager handle context for frozen agents?
Exploring whether a separate trained system can effectively manage a frozen agent's context window. This matters because many deployed agents are closed-source and can't be retrained, yet they suffer from context degradation.
contrasts a separate trained compression manager (AdaCoM) with AutoCompact's single jointly-trained policy
-
Which coding harness components matter most in different conditions?
Can individual harness components—planning, context management, action space—be evaluated separately rather than as a package? This matters because practitioners need to know which components to prioritize given their constraints.
AutoCompact's own budget ablation (19.9% gain at $0.10 vs 1.9% at $4.00) corroborates that context management pays off most under tight budgets
-
Does agent memory degrade when continuously consolidated?
Can consolidating agent experiences into summaries actually harm long-term performance? Research on ARC-AGI tasks suggests continuous memory updates may reduce capability below the no-memory baseline.
both find that compaction or consolidation quality, not mere presence, determines whether it helps or hurts task performance
-
Why does self-correction training on offline data fail?
Can language models learn to correct their own mistakes through supervised training on correction examples? This explores whether distribution mismatch and behavior collapse prevent self-correction from emerging.
evidence for — SCoRe's distribution-mismatch explanation for SFT failure matches why SFT imitation of summaries underperforms RL in A
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
- Can Large Reasoning Models Self-Train?
- Training Language Models to Self-Correct via Reinforcement Learning
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
- On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
- A Survey on Post-training of Large Language Models
- Self-Rewarding Language Models
Original note title
AutoCompact's task-success RL reward teaches summary self-consistency that SFT imitation of compaction summaries does not