PaperScope
LIVE · 2026-09-09 05:40 UTC

On BatchNorm Forward Modes in Value-Based Reinforcement Learning

Daniel Palenicek, Mikael Henaff, Scott Fujimoto, Koustuv Sinha

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06421 v1
Category
Submitted
2026-09-06

Abstract

Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch identified by CrossQ. We show for target-based C51 and target-free PQN that the simple choice between running and batch statistics at specific forward passes can reverse this degradation. In C51, switching the BN bootstrap forward to batch-statistic mode significantly improves performance over unnormalized and LayerNorm baselines and scales stably with update-to-data ratios up to 12. In PQN, using batch-statistics for both action selection and bootstrapping recovers performance from the failing running-statistic configuration. Across 26 Atari games at 400M frames, this configuration achieves a higher final aggregate score than PQN with LayerNorm. Our results show that carefully configured BN can substantially improve discrete-action value learning, and that its forward protocols are an essential part of the algorithm specification.

arXiv abs page · PDF · same-day batch