PaperScope
LIVE · 2026-10-06 05:40 UTC

What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects

Bowen Xu, Boyu Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05533 v1
Category
Submitted
2026-10-04

Abstract

Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation. We preregistered a study of where these gains come from, with three small models, three benchmarks, replication and test partitions, a GEPA search arm and six evaluation defects injected one at a time, and we report all 47 primary endpoints. Turning thinking off raised accuracy over a capped thinking setting in 5 of 9 model-benchmark cells, and in each the gain came mostly from questions where the capped setting gave no readable answer. The thinking-off setting was not meaningfully worse than a rescue configuration or four GEPA-selected harnesses in 11 of 13 comparisons, and lost to the rescue on GSM8K for two models. GEPA repaired its broken starting points, but none of its selected harnesses was more accurate than the thinking-off setting. A thinking budget in the serving engine, which also allows a longer answer, lowered truncation and raised the parse rate in 6 of 9 cells. In 6 of 15 evaluable defect-model pairs, replication through the same pipeline reproduced the defect's distortion instead of revealing it. On the LongevityBench multiple-choice tasks, only the longevity-tuned model beat the strongest constant-label baseline.

Comment: 43 pages, 2 figures, 35 tables. Ancillary files in anc/: the frozen preregistration, its addenda (row keys of five rows withheld) and the aggregate analysis report, with a README

arXiv abs page · PDF · same-day batch