Recursive Self-Rewrite turns guided terminal runs into standalone capability

Researchers have posted a preprint (arXiv:2610.02826) describing Recursive Self-Rewrite, a framework that lets a single base model — Qwen-3.8-27B — learn from successful solution trajectories produced with dedicated helper tools and rewrite them into training paths that work in a general environment without external intervention. The problem is familiar: when a model solves complex terminal tasks with the aid of "harnesses," special scripts that guide, check or correct it, the acquired knowledge remains tied to those harnesses and does not transfer to deployment. RSR is designed to sever that dependency.
The system runs three components in a loop. A planner extracts a runbook — a structured action sequence — from a successful trajectory. A critic filters out leaks of the solution or the verifier and steers recursive correction. An executor runs the approved runbooks in clean sandboxes. The pipeline converted 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised fine-tuning (SFT). Three different harnesses operated together on roughly 3,000 independently developed terminal tasks and jointly solved 759 of them, 34.3% more than the strongest harness alone.
Benchmark results show consistent improvement but large gaps remain. Training on the rewritten trajectories lifted pass@3 on Terminal-Bench 2 from 57.0% to 74.2%, and on Terminal-Bench Hard (developed by the authors) from 39.0% to 63.0%. On the harder end, Terminal-Bench 4 rose only from 1.5% to 9.1%, and Software Terminal-Bench from 3.0% to 6.0%. Process reward on Long-Horizon Terminal-Bench climbed from 0.21 to 0.29. The numbers confirm the method works, yet long-horizon, complex tasks remain a significant challenge even after the data was multiplied and diversified.
Several pieces are missing from the picture. The paper does not report comparisons against larger models or against other synthesis methods — distillation from a stronger model, for instance. The hardest benchmarks are the authors' own, making independent evaluation difficult. This is a preprint that has not yet undergone peer review; as of now no models, datasets or Spaces citing it appear on Hugging Face. The distinction between "open weights" and "open source" matters here: Qwen-3.8-27B is available under open weights, but the RSR code itself has not been released.
The bottom line: RSR demonstrates how experience gathered under "protected" conditions can be turned into capability that works in a clean environment, a practical step toward agents that function without close supervision. The gains on Terminal-Bench 2 and Hard are impressive for a 27B-parameter base model, but the gaps on Terminal-Bench 4 and Software Terminal-Bench underscore that long-horizon software tasks are still far from solved. Worth watching for a code release and for whether the results hold up on external benchmarks.