PaperScope
LIVE · 2026-09-29 05:40 UTC

Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

Christian Moya, Elliott Thornley, Guang Lin

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35677 v1
Category
Submitted
2026-09-28

Abstract

In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.

arXiv abs page · PDF · same-day batch