PaperScope
LIVE · 2026-09-09 05:40 UTC

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

Mahmoud Selim, Cristina Cipriani, Karl H. Johansson

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06882 v1
Category
Submitted
2026-09-07

Abstract

Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion latents through the distribution of executed actions induced by the denoising process. We show that this construction admits a precise semantic interpretation and derive a noisy-space policy gradient (NSPG) that optimizes noisy latents using only clean action-space value estimates. Building on this result, we formulate a KL-regularized policy improvement over noisy latents and show that the resulting objective admits a diffusion-compatible regression form, avoiding backpropagation through the denoising process. Empirical results on state-based D4RL benchmarks and vision-based OGBench tasks demonstrate that the proposed noisy-space objective provides a principled and effective basis for training diffusion policies in offline reinforcement learning. Project webpage: https://mahmoud-selim.github.io/NSPG/

Comment: Accepted at the 43rd International Conference on Machine Learning (ICML 2026)

arXiv abs page · PDF · same-day batch