PaperScope
LIVE · 2026-09-29 05:40 UTC

Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles

Yongyi Guo, Zifan Xu, Ziping Xu, Kelly W. Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32170 v1
Category
Submitted
2026-09-26

Abstract

The performance of online reinforcement learning depends critically on design choices, especially those that affect exploration. These choices are often selected by fitting a simulator to offline data, evaluating candidate algorithms in that simulator, and deploying the best-performing one. The simplest Plug-In selection rule simply selects the best performing algorithm on the fitted simulator, making evaluations unreliable when the offline data used to fit the simulator are limited. We investigate Uncertainty-Aware selection, which forms an ensemble of simulators---for example, obtained by bootstrap resampling---and selects the online algorithm with the best average performance across the ensemble. While ensemble-based approaches have been used to mitigate distribution shift and facilitate sim-to-real transfer, we formally show that this approach can mitigate the effects of limited data when fitting the simulator and theoretically has significant regret gains compared to Plug-In selection in multi-armed bandits. We also empirically investigate the Uncertainty-Aware selection approach in deep RL experiments on robotic control tasks that involve selecting reward-shaping hyperparameters, and show that it leads to more reliable selection and improved online performance.

arXiv abs page · PDF · same-day batch