When Can Old Evaluations Certify a New Model? Label-Efficient Release Decisions under Evaluator Drift
Joyanta Jyoti Mondal, Mridul Banik, Md. Shifatul Ahsan Apurba, Md Masud Al Mahmud
Abstract
Releasing a model update requires certifying that its current-population risk stays below a threshold. Trusted labels are expensive, while a cheap evaluator, such as an LLM judge, scores every example. Reusing evaluator errors from earlier audits is tempting, but when may such evidence replace current labels? It depends on the status of history. If the errors can change invisibly, no label-free test detects the change, and every valid, useful certifier must keep buying labels at a rate we characterize; if a bound on the change is assumed, label-free certification is valid at an explicit error cost. For the middle ground, where history is informative but untrusted, we propose \emph{portfolio vigilance}, a sequential certifier mixing a betting expert guided by history with one that learns only from current labels; history affects only how it bets, so validity holds for any history. The contribution is not prior-informed betting or expert mixtures, but separating history that may enter validity from history that may only guide label collection. In a canonical model, accurate history shortens decisions but never raises the evidence growth rate; stale history can destroy it. On held-out CIFAR-10N and DICES-990 data, portfolio vigilance needs 0.465 (95\% CI $[0.327,0.575]$) and 0.740 ($[0.618,0.877]$) times the labels of a matched prediction-powered monitor, with no observed false certification, and fewer labels on all six external blocks. Under corrupted advice it stays within 8.0\% of its better component, while trusting history alone costs up to 1.66 times as much. In post-confirmatory repeated-judge experiments on DICES-990 and ToxicChat, changing a fixed LLM judge's rubric moves its scores beyond run-to-run variation; the portfolio then needs 0.790 ($[0.667,0.909]$) and 0.631 ($[0.520,0.770]$) times the labels of the matched monitor, and fewer than trusting history alone.