Generalizability of harness evolution versus repeated sampling
Determine whether automatic harness evolution for large language model agents yields generalizable improvements in harness design, or whether the observed gains are primarily due to repeated sampling during the search process under matched feedback and inference budgets.
References
The existing evaluation protocol therefore leaves a fundamental question unresolved: Does harness evolution yield generalizable improvements in harness design, or are its gains primarily due to repeated sampling?
— Rethinking the Evaluation of Harness Evolution for Agents
(2607.12227 - Wang et al., 14 Jul 2026) in Section 1 (Introduction)