Overfitting of Spectral Gradient Descent: How Matrix Geometry shapes Generalization and Implicit Bias
Guillaume Braun, Ichiro Hashimoto, Masaaki Imaizumi
Abstract
We study the generalization of spectral gradient descent (SpecGD) in overparameterized matrix classification with corrupted labels. Each input combines a shared low-rank signal with a rank-one sample-specific perturbation, referred to as a shortcut, that enables memorization but does not generalize. We contrast collapsed shortcuts, which share a singular direction, with dispersed shortcuts, which occupy distinct singular directions. Changing only this geometry can reverse the relative generalization of GD and SpecGD: collapsed shortcuts can favor SpecGD, while dispersed shortcuts can favor GD. In the dispersed regime, exact shortcut orthogonality eliminates the signal from the late-stage SpecGD direction, while vanishing random correlations collectively generate a small but generalization-relevant signal through a second-order effect. To identify the direction selected by SpecGD, which the spectral max-margin problem alone does not determine, we combine a refined analysis of its dual with the exponentiated-gradient dynamics of normalized loss weights. Finally, we show that a single SpecGD step can already interpolate and generalize well, while continued training converges to a direction with substantially worse generalization.