Command Palette
Search for a command to run...
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Zongxia Li Yucheng Shi Zhongzhi Li Junyao Yang Ruhan Wang Chengsong Huang Fuxiao Liu Haitao Mi Jordan Boyd-Graber Leowei Liang
Abstract
Successful trajectories on dificult tasks are valuable supervision for model improvement. Diferent harnesses enable the same model to solve these tasks in diferent ways. Their successful trajectories provide useful experience for self-improvement, but they also contain controller interventions and workflow conventions that may be unavailable under a general harness. We propose Recursive Self-Rewrite (RSR), a framework that converts these experiences into reusable model capabilities through recursive trajectory self-rewrite. We use a single base model, Qwen-3.8-27B to discover successful solutions under diverse harnesses and rewrite them for learning under a general harness. RSR contains a planner that extracts useful procedures into runbooks, a critic that screens for verifier and solution leakage, and rejects candidates and recursively regenerate using critique feedback, and an executor that follows the qualified runbooks to solve each task in a fresh sandbox under a general harness. We show that using three harnesses expands the range of task domains that Qwen-3.8-27B can solve and yields more valuable successful trajectories; with RSR, we further reconstruct these into a larger set of high-quality trajectories. We collect experience from approximately 3K self-curated terminal tasks. Across 3K tasks, the union of the three harnesses solves 759 tasks, which is 34.3% more than the strongest individual harness in the recorded pool. We further use RSR to rewrite successful source trajectories, expanding the training set from 2,001 to 11,094 high-quality trajectories for finetuning Qwen-3.8-27B. Training it on these rewritten trajectories outperforms both the base model and direct trajectory SFT: pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. On Long-Horizon Terminal-Bench, process reward rises from 0.21 to 0.29.
One-sentence Summary
Recursive Self-Rewrite (RSR), proposed by researchers from Tencent HY LLM Frontier, the University of Maryland, College Park, and other institutions, recursively rewrites successful trajectories from diverse harnesses into reusable runbooks via a planner, a leakage-screening critic, and an executor, expanding Qwen-3.8-27B finetuning data from 2,001 to 11,094 trajectories and improving Terminal-Bench 2 pass@3 from 57.0% to 74.2%.
Key Contributions
- The paper introduces Recursive Self-Rewrite, a framework that converts successful trajectories from diverse specialized harnesses into verified demonstrations under a general harness. It consists of a planner that extracts reusable runbooks, a critic that screens for verifier and solution leakage and recursively regenerates rejected candidates, and an executor that follows qualified runbooks in fresh sandboxes.
- It shows that using multiple harnesses as discovery tools broadens task coverage: across approximately 3,000 self-curated terminal tasks, the union of three harnesses solves 759 tasks, 34.3% more than the strongest individual harness.
- It demonstrates that rewriting these trajectories expands the training set from 2,001 to 11,094 high-quality examples and improves fine-tuned Qwen-3.8-27B over the base model and direct trajectory SFT. Pass@3 rises from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on Terminal-Bench Hard, and from 3.0% to 6.0% on Software Terminal-Bench, while process reward rises from 0.21 to 0.29 on Long-Horizon Terminal-Bench.
Introduction
The authors study self-improvement for AI agents on difficult terminal tasks, where performance depends not only on the model but also on its execution harness, the system that controls observations, tool use, verification, recovery, and stopping. Prior work shows that specialized harnesses can improve task success without changing model weights, and that different harnesses solve complementary subsets of tasks. However, trajectories collected from these harnesses mix harness-specific controller interventions, prompts, workflow logic, and stopping rules with the underlying problem-solving behavior. Naively training on such data can cause the model to depend on external control patterns that are unavailable under a general harness at inference time. To address this, the authors propose Recursive Self-Rewrite, which collects successful multi-harness trajectories and rewrites them into verified trajectories under a general harness using the same base model in planner, critic, and executor roles. This enables supervised finetuning from the model's own harness-assisted experience and improves Qwen-3.8-27B across several terminal-task benchmarks.
Dataset
The authors describe the terminal-task dataset as follows:
- Composition and sources: The training task pool is built from two sources. SWR is a self-constructed collection of about 2,500 terminal tasks across 50 domains. The authors also include 420 filtered and modified tasks from the RST dataset.
- Domain coverage: SWR spans software usage, biology, chemistry, physics, hardware, operations, and security.
- Scale: Together, these sources define a pool of roughly 3,000 tasks.
- Filtering and modification: The RST subset is filtered and modified to have increased difficulty.
- Usage: The pool is used for terminal-task training, where the model must select tools, reason about environment feedback, and carry out task-specific procedures.
- Additional processing details: The provided text does not specify schema, cropping, metadata construction, or mixture ratios.
Method
The authors propose Recursive Self-Rewrite, a method designed to improve a model under a general-purpose harness by learning from successful trajectories discovered under diverse specialized harnesses. The core idea is that different harnesses help the same base model solve different difficult tasks, and those successful solutions can be rewritten into demonstrations compatible with a single general harness. The overall process is illustrated in the framework diagram below.
The method operates through three primary stages: Multi-Harness Discovery, Trajectory Rewriting, and Verification.
Multi-Harness Discovery The authors utilize multiple discovery harnesses that differ in how they provide control and support during execution, such as progress tracking, continuation, validation, state management, or recovery from failure. Because different harnesses can be effective on different types of tasks, they expand the range of successful solutions beyond what the model could achieve with just one harness. The base model is run under these multiple discovery harnesses to collect successful trajectories.
Trajectory Rewriting for Experience Learning Successful trajectories collected under different harnesses record how the model solves tasks with different workflow logic. The goal is to transform these successful solutions into learning experiences that the model can practice and learn from under a general harness. Trajectory rewriting involves three distinct model roles: a planner, a critic, and an executor.
- Planner: The planner reconstructs a runbook for each successful source trajectory, providing a structured description of how the task was solved. Before planning, the source trajectory is compacted by retaining the task instruction, the model’s actions, and the environment’s observations, while removing harness-specific control messages. The runbook summarizes the required end state, key milestones, useful checks, recovery strategies, and common pitfalls. To reduce direct answer transfer, runbooks describe the task, relevant interfaces, and validation procedures without directly providing the finished deliverable. Multiple runbook candidates are sampled for each source trajectory.
- Critic: The model itself acts as a critic to filter candidate runbooks before they are used for execution. Deterministic checks are first applied, such as schema validation, removal of known artifacts, and rejection of unsupported tool references. Then, a model-based critic that sees only the public task instruction and the candidate runbook determines whether the runbook provides useful procedure or leaks information that the executor should not receive. Only runbooks that pass this screening stage are retained for rewriting.
- Executor: For each approved runbook, the executor re-solves the task in a fresh sandbox under the general harness. The runbook is provided as private guidance during generation but is never written into the public trajectory. The executor must therefore produce a new trajectory based on the current environment rather than replaying the source trajectory.
Trajectory Filtering and Verification The model itself is used to flag values that appear in a trajectory but cannot be derived from the task or the environment, which indicates hidden answer transfer. Demonstrations containing such values are discarded. For finetuning, only the public interaction history, comprising the task instruction, environment observations, and the model’s responses, is kept, while the runbook and critic conversation are removed. At inference time, the finetuned model runs under the general-purpose harness alone, without the source harnesses or private runbooks, ensuring that any planning, checking, recovery, or continuation comes from the model itself.
Case Study Examples The authors examine rewritten tasks to demonstrate how a source experience can be transformed into new trajectories under the general harness. As shown in the figure below, the examples illustrate a passing and a failing execution guided by the same runbook in each case, alongside representative command excerpts from the source and both rewrites.
Experiment
The experiments evaluate a self-improvement pipeline that collects successful terminal-task rollouts from three complementary execution harnesses, Terminus 2, StateM, and Recursive Self-Reflect Terminus, then reconstructs those experiences into standardized trajectories under the general harness. They show that different harnesses unlock different problem-solving behaviors and jointly expand task coverage, while direct fine-tuning on source rollouts produces mixed gains and can introduce looping failures. Rewriting successful trajectories through recursive self-rewrite leads to cleaner, more generalizable training data and improves performance over both the base model and direct fine-tuning, with additional partial progress on long-horizon tasks.
Combining rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM expands solved-task coverage beyond any single harness or pair. The three-harness union solves the most tasks overall and adds a substantial number over the strongest individual harness, with relative gains on both RST and SWR. Individual harnesses show complementary strengths, as leaders vary across benchmark splits. The three-harness union yields the highest overall task coverage and a roughly one-third relative increase over the strongest individual harness. Across individual harnesses, Recursive Self-Reflect Terminus leads on RST, while Terminus 2 leads on SWR and the pooled set; pairwise unions consistently exceed any single harness.
Combining rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM broadens discovery coverage beyond any single harness. The pooled collection contains 2,001 successful trajectories covering 759 distinct tasks, which is more solved tasks than the strongest individual harness. Per-rollout success rates are similar across harnesses, indicating that multi-harness pooling mainly increases task coverage. The strongest individual harness solves 565 tasks, while the pooled union solves 759 tasks, adding 194 solved tasks and a relative coverage increase of 34.3%. Recursive Self-Reflect Terminus achieves the highest individual rollout success rate at 15.9%, StateM has the lowest at 11.7%, and the pooled set reaches 13.7%.
Across passing trajectories, the three harnesses show distinct execution styles. Terminus 2 produces shorter trajectories with more exploration-oriented commands and more passing rollouts, while StateM produces longer trajectories with more commands per turn but less exploration. RSRT sits between them on trajectory length and exploration, with the most completion claims per trajectory and a small share of passes occurring after rejection. Terminus 2 has the largest number of passing trajectories and the shortest average and median turn counts among the three harnesses. StateM shows the longest median trajectory length and highest commands per turn, but the lowest exploration command share. RSRT records the most completion claims per trajectory and is the only harness with a notable percentage of passes after rejection.
Rewrites guided by the same runbook tend to have more similar command-level behavior than rewrites using different runbooks. This pattern is stronger for Markdown, while OpenFOAM shows a weaker effect for command metrics. Workflow-level similarity is less consistent and varies by task. Same-runbook pairs show higher exact command overlap and command-order similarity than different-runbook pairs, with a clearer gap for Markdown. Tool-transition and action-sequence similarity are not consistently higher within runbooks; OpenFOAM shows little tool-transition difference and slightly lower within-runbook action-sequence similarity.
RSR outperforms both the base model and Direct SFT across the reported benchmarks, with higher pass@3, higher mean per-run pass rate, and better process reward on LHTB. Direct SFT shows mixed results, improving on several benchmarks but declining on TB2, where training without rewriting can introduce looping behaviors. The results suggest that scaling and standardizing successful experiences under a general harness strengthens generalization. RSR achieves the highest pass@3 and mean per-run pass rate across all five reported benchmarks. Direct SFT has mixed results, improving over the base model on TBH, TB3, and TB4 but declining on TB2 and falling below its base mean per-run pass rate.
The experiments evaluate multi-harness rollouts from Terminus 2, Recursive Self-Reflect Terminus, and StateM, along with runbook-guided rewrite similarity and a comparison of RSR against base and Direct SFT models. Combining the three harnesses yields complementary task coverage and the largest set of solved tasks, while the harnesses show distinct execution styles and same-runbook rewrites are more similar at the command level, especially for Markdown. RSR consistently outperforms the base model and Direct SFT, whereas Direct SFT improves on several benchmarks but can degrade on others through looping behaviors.