Automated tool reproduces harness bugs in agents and builds live benchmark of 200 cases
Researchers have posted AgentBug-Smith to arXiv, an automated approach for discovering and reproducing harness bugs in open-source agentic systems. Existing benchmarks hold a small, fixed set of bugs and require hundreds of human hours to construct; the new tool runs continuously against live repositories. Harness bugs affect the agent's own execution and testing infrastructure rather than business logic, making them especially difficult for state-of-the-art software agents to fix.
Across several backbone LLMs, AgentBug-Smith achieved success rates 10.67% to 27.56% higher than general-purpose bug-reproduction techniques designed for conventional software. The gap shows that the distinctive traits of harness bugs — environment dependencies, non-deterministic execution state, and brittle internal interfaces — demand a dedicated methodology; generic tools simply do not close it.
Using the tool, the team built Live-Harness-Bench, a live, extensible benchmark that currently contains 200 fully reproduced harness bugs. The main advantage is not just volume but dynamism: every new failure in an open-source agent can automatically become an executable evaluation instance, without waiting for a human team to characterize, isolate, and document the bug manually.
The researchers demonstrated two downstream uses. First, the benchmark enabled systematic evaluation of leading software agents, revealing limited ability to fix real harness bugs. Second, the corpus of fixes served as a knowledge base from which repair skills were distilled for reuse; integrating those skills into existing agents lifted the harness-bug fix rate by 6.32%. The figure is modest, but it proves that an automated feedback loop — failure, reproduction, repair, distillation, improvement — is realizable.
The work, which has not yet undergone peer review, frames the dual infrastructure of reproduction tool and live benchmark as a foundation for continuous evaluation and improvement of agents on the harness-bug repair task. The stated goal is to enable recursively self-improving agents by turning real-world failures into actionable knowledge. For now, the code and data are public, and the benchmark is open to community expansion.