PaperScope
LIVE · 2026-09-23 05:40 UTC

Calibration Count Reuse: Validity Does Not Determine Efficiency

Rudra Chopra

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.25138 v1
Submitted
2026-09-21

Abstract

Calibration count reuse raises separate validity and efficiency questions. We give a validity criterion for general count-dependent nonconformity scores: transferring one count from another class to the scored class must not improve its conformity. A leave-self-out full conformal reference proves the criterion without requiring normalization or preservation of same-class score order. For a common separable transformation, universal exchangeable validity is equivalent to being nondecreasing in the count, provided $Kα\geq 1$; normalized multiplicative weights obey the complementary nonincreasing condition. Additive penalties are covered under the stated information restrictions. Efficiency has no parallel ordering: two iid constructions make the same valid rule improve or worsen expected size at unchanged coverage. An expanded 55-rule study finds no resolved advantage from selected live-count rules over uniform weights. Image studies identify undercoverage under iid resampling, including at numerical convergence. Separately, execution of the released Conf-OT pipeline on its DTD and Aircraft benchmark subsets produces near-nominal median coverage under fixed stratified counts. The native results are reported separately from the iid analyses, without treating a benchmark observation as a universal guarantee. The findings separate validity, classifier confidence, numerical convergence, and population-specific efficiency.

Comment: 33 pages, including appendices

arXiv abs page · PDF · same-day batch