PaperScope
LIVE · 2026-09-29 05:40 UTC

When the Score Becomes the Target: Rethinking Metric Validity in Autonomous Driving

Morui Zhu, Deyuan Qu, Qi Chen, Kentaro Oguchi, Qing Yang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34440 v1
Category
Submitted
2026-09-28

Abstract

Driving benchmark scores are increasingly used not only for evaluation but also as optimization targets. This raises a fundamental question: do score gains remain reliable evidence of driving improvement once the score itself is optimized? We address this question by examining how the scoring process responds to changes in driving behavior and whether the resulting gains persist under repeated execution and replanning. We decompose the process into execution, measurement, subscore mapping, and aggregation. Controlled interventions reveal substantial behavioral changes that receive little score response because distinctions are omitted, thresholded, or attenuated between requested and executed motion. Closed-loop comparisons further show that optimization gains can reverse when the execution interface changes, demonstrating their dependence on how requests are executed and returned as feedback. Together, these findings connect the behavioral distinctions preserved by a metric to the conditions under which its gains transfer. Metric validity under optimization therefore requires examining both what the scoring process measures and how the optimized behavior is executed.

Comment: 28 pages including supplementary materials

arXiv abs page · PDF · same-day batch