PaperScope
LIVE · 2026-09-30 05:40 UTC

Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks

Zhaojun Peng

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36859 v1
Category
Submitted
2026-09-29

Abstract

We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-γP_π)^{-1}$ can provide the multiplier stability that classical nonconvex ADMM analyses often obtain from a designated smooth block. Starting from this mechanism, we establish convergence under controlled Markov sampling and then under stochastic observations using an empirical Bellman surrogate that jointly represents the random residual and its Jacobian. Markov mixing, initialization drift, observation noise, and decaying bias enter as one operator perturbation, avoiding unbiased product and double sampling requirements. When the perturbations are square summable, the true KKT residual converges almost surely to zero. Under a finite conditional fourth moment condition, a companion iterate satisfies $ \mathbb{E}[\widetilde G_{K+1}] \le A/T+(B/T)\sum_{k<T}m_k^{-1}, $ which becomes $O(T^{-1}+T/N)$ for total Markov sample budget $N$, giving $O(ε^{-1})$ iteration complexity and $O(ε^{-2})$ sample complexity for squared KKT accuracy $ε$. Beyond stationarity, discounted occupancy coverage yields $J^\star-J(π)=O(\sqrt G)$ for direct tabular policies, so covered exact KKT points are globally optimal, while a statewise quadratic Bellman-improvement condition sharpens the relation to $O(G)$. Finally, nonlinear policy, projected Bellman, and explicit occupancy formulations exhibit the same chain of operator invertibility, dual representation, and multiplier stability. This supports discounted operator invertibility as a reusable structural principle for primal-dual reinforcement learning.

arXiv abs page · PDF · same-day batch