PaperScope
LIVE · 2026-10-09 05:40 UTC

When KL Regularization Misfires in Group Policy Optimization

Fei Ding

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.12161 v1
Category
Submitted
2026-10-08

Abstract

Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.

Comment: 27 pages

arXiv abs page · PDF · same-day batch