Clipped or Unclipped? Finite-Sample Trade-offs for Averaged SGD under Heavy-Tailed Noise
Alexandra Suvorikova, Egor Gladin, Darina Dvinskikh, Artem Agafonov, Mohammad Alkousa, Yuriy Dorn, Vladislav Matyukhin, Alexander Gasnikov
Abstract
Gradient clipping is widely used to stabilize training, but it need not improve the statistical accuracy of averaged SGD, even under heavy-tailed noise. We derive a finite-sample comparison of clipped and unclipped Polyak-Ruppert averaged SGD under finite conditional $p$-th moments, $p\ge2$. Our main result gives explicit accuracy and confidence conditions under which, for $p>2$, the Gaussian term dominates the unclipped deviation bound, so clipping need not improve its leading order. By balancing clipping bias and concentration, we obtain a bound in which the heavy-tail correction depends logarithmically rather than polynomially on the inverse failure probability. At $p=2$, this improves the confidence dependence of the leading bound. We establish sharpness of the unclipped heavy-tail term through an exact one-dimensional quadratic recursion and extend the comparison to projected convex SGD. We also prove concrete costs of clipping: every fixed finite threshold increases asymptotic variance on a scalar Gaussian quadratic, while whole-gradient clipping can shift the limiting point under asymmetric noise.