PaperScope
LIVE · 2026-09-29 05:40 UTC

Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations

Amitakshar Biswas, Yuhan Li, Ruoqing Zhu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33186 v1
Category
Submitted
2026-09-27

Abstract

Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing weight. In ideal settings, this mixed Bellman formulation can provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting. To solve this mixed residual optimization, we adopt a minimax formulation involving a critic function. Unlike standard approaches that rely on a fixed functional class, we construct a data-dependent critic representation using predicted future feature directions which effectively induces a kernel adapted to the underlying transition dynamics. This allows the critic to focus on directions that are most relevant for the estimated Bellman error. To control overfitting, we use sample splitting to construct the critic and estimate the value function on separate data subsets. Simulation studies and MetaWorld tasks illustrate the effect of the mixing parameter and show that intermediate residual combinations can improve value estimation in challenging settings.

arXiv abs page · PDF · same-day batch