PaperScope
LIVE · 2026-10-07 05:40 UTC

A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral

Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.07084 v1
Category
Submitted
2026-10-05

Abstract

An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert quality and an unfinished correction: early errors can discourage queries that would be valuable after learning. We propose ORUCB, which pools shared and expert-specific polynomial responses. A bound on cumulative response-learning error calibrates confidence-weighted risk regression and exploration, allowing the router to account for this error when deciding which answers to buy. Under bounded residuals and disagreements, a fixed feasible model of optimal responses, and linear models of free and optimal queried risk, the calibrated algorithm achieves high-probability pseudo-regret $O(\sqrt T\log(T+1))$ over $T$ rounds for fixed problem parameters. The guarantee permits singular answer distributions and misspecified shared responses; optimality is relative to the bounded response class. On four test streams, the selected cubic policy has lower fee-inclusive cost than seven baselines that deploy answers unchanged. Comparisons with a common correction learner examine routing, while six-price comparisons measure cost and query rates.

arXiv abs page · PDF · same-day batch