PaperScope
LIVE · 2026-10-09 05:40 UTC

Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation

Biao Xiang, Ali Eshragh, Yuexing Li, Kai Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10974 v1
Submitted
2026-10-07

Abstract

Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.

Comment: Accepted to NeurIPS 2026. Code available at https://github.com/Lygist/Budgeted-Multi-Source-Counterfactual-Annotation-for-Off-Policy-Evaluation

arXiv abs page · PDF · same-day batch