Generalizability of harness evolution versus repeated sampling

Determine whether automatic harness evolution for large language model agents yields generalizable improvements in harness design, or whether the observed gains are primarily due to repeated sampling during the search process under matched feedback and inference budgets.

Background

The paper critiques common evaluation protocols for automatic harness evolution, where harness search and final evaluation are conducted on the same benchmark, potentially conflating genuine harness improvements with gains from additional test-time computation such as repeated sampling.

To disentangle these effects, the authors highlight the need to compare harness evolution with test-time scaling baselines under matched feedback and inference budgets, raising the question of whether harness evolution produces reusable, generalizable harness improvements beyond benefits from repeated sampling.

References

The existing evaluation protocol therefore leaves a fundamental question unresolved: Does harness evolution yield generalizable improvements in harness design, or are its gains primarily due to repeated sampling?

Rethinking the Evaluation of Harness Evolution for Agents  (2607.12227 - Wang et al., 14 Jul 2026) in Section 1 (Introduction)