PaperScope
LIVE · 2026-09-09 05:40 UTC

Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks

Yiming Ying

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06430 v1
Category
Submitted
2026-09-06

Abstract

We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss $(-yu^\top x)_+$. This representation makes a stability analysis possible. We derive an exact distance identity for two coupled updates and prove approximate non-expansiveness of the common-example map, with a quadratic defect only when the two latent margins straddle zero. We then obtain explicit $\ell_2$ on-average model-stability and generalization bounds, transferring stability isometrically from the latent vector to the full first-layer matrix. Combining stability with a standard optimization bound yields an explicit excess induced-risk guarantee and the rate $O(n^{-1/2})$ when $T=n^2$. Under margin separability, a complementary argument gives the optimal-order $O(R^2/(γ^2n))$ expected excess misclassification error for a randomized one-pass STE iterate and a corresponding majority-vote bound.

arXiv abs page · PDF · same-day batch