PaperScope
LIVE · 2026-09-03 05:40 UTC

Independent Reinforcement Learning in Discounted Markov Games

Asrin Efe Yorulmaz, Ugur Aydin, Tamer Basar

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.00504 v1
Submitted
2026-09-01

Abstract

In this work, we study radically uncoupled learning in discounted general-sum Markov games. Assuming ``$\mathsf{ETH}$ for $\mathsf{PPAD}$", we show that, for every fixed discount factor, there is no polynomial-time algorithm for computing inverse-polynomially accurate coarse correlated equilibria in discounted general-sum Markov games when players learn independently in decentralized settings. Complementing this hardness result, we provide what appears to be the first \emph{radically uncoupled} algorithm with sub-exponential convergence guarantees to coarse correlated equilibria in discounted general-sum Markov games without imposing any structural restrictions on the game. Our algorithm is a \emph{layered} variant of optimistic mirror descent with an increasing step-size schedule tailored to the multi-agent setting. Finally, we develop both full-feedback and partial feedback versions of the aforementioned algorithm and establish sub-exponential convergence guarantees for each case.

Comment: 54 pages, 3 figures

arXiv abs page · PDF · same-day batch