PaperScope
LIVE · 2026-09-29 05:40 UTC

Deep Epistemic Value Functions for Optimistic Exploration

Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35525 v1
Category
Submitted
2026-09-28

Abstract

Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.

arXiv abs page · PDF · same-day batch