PaperScope
LIVE · 2026-09-15 05:40 UTC

When Compliance Data Masquerades as Evaluation: Measurement Validity for Deployed AI Systems

Hung-Yu Lin, Xingran Huang, Qiming Guo, Jinwen Tang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.13642 v1
Category
Submitted
2026-09-12

Abstract

We argue that a recurring failure in the evaluation of deployed AI systems occurs when data collected for operational monitoring or regulatory compliance are interpreted as if they were designed for comparative evaluation. Automated driving provides a concrete example of this problem. U.S. disengagement and crash-reporting regimes produce valuable operational evidence, but differences in reporting scope, exposure, deployment domain, event capture, and comparator construction limit the safety claims that can be supported from these measurements alone. We frame this issue as a measurement-validity problem in AI evaluation rather than as a transportation-specific data limitation. We argue that comparative claims about deployed AI systems require alignment between the intended capability, measured outcome, exposure opportunity, deployment domain, data-generation process, and evaluation comparator. Using automated-driving safety evaluation as a case study, we propose an evaluation contract that makes these assumptions explicit before operational data are interpreted as evidence of comparative performance. The broader implication is that data useful for monitoring deployed AI systems are not automatically valid benchmarks for evaluating them.

arXiv abs page · PDF · same-day batch