Researchers propose branching architecture that lets agent harnesses improve without getting stuck
in local optima A new arXiv pre-print introduces Mixture of Self-Improving Branches, a method that makes the self-improvement process of harnesses adaptive. Instead of a single search trajectory with a fixed development set and proposal policy — the approach Meta-Harness takes — the system spawns parallel branches, each with a varying development subset and a proposal policy updated from its own search history. The result is a set of complementary harnesses that a router selects between at runtime, without access to test results.
The single-path problem
Meta-Harness implements recursive self-improvement (RSI) through successive generations of code generation and evaluation, but keeps the development set and proposal policy static. That constraint forces evolution along a single path, increasing the risk of converging to a local optimum — exactly the problem the researchers set out to solve.
How the branches work
Each branch retains cases solved by more of its own top harnesses than by top harnesses in other branches, discards cases solved by all top harnesses across all branches, and updates its proposal policy based on its internal search history. The mechanism produces diverse optimization directions instead of a single narrow push.
Router without test-time leakage
To deploy the complementary harnesses, the authors propose a router that picks one branch head per new input before execution. Both the selection and the router configuration rely solely on development data. There is no learning from the test set, a point the authors emphasize explicitly.
Results: relative gains on three benchmarks
On Olympiad-level mathematical reasoning the system achieves a 34.8% relative improvement over Meta-Harness; on Terminal-Bench 2.0, 11.6%; and on SWE-bench Lite, 3.8%. The numbers come from a direct comparison under the same setup, with harness selection and router calibration performed only on development data.
Caveats: pre-print, not peer-reviewed
The paper appears as a pre-print on arXiv and has not yet undergone peer review. Code and data are available, but metrics were measured under lab conditions on defined samples; there is no indication yet of how the approach behaves in production or in more heterogeneous environments.