PaperScope
LIVE · 2026-10-06 05:40 UTC

Underscoring the Problem: Why Softpick Fails at Initialization

Aryan Sood, Jaikaran Singh, Ishaan Bansal

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05488 v1
Category
Submitted
2026-10-04

Abstract

Softmax attention gives every token a nonzero weight, which in trained models concentrates into attention sinks and massive activations that widen the dynamic range low-precision inference must cover. Softpick removes this constraint by rectifying scores, eliminating sinks and lowering hidden-state kurtosis, but its advantage fades at scale. We reframe this failure as a normalization problem. Softpick's denominator splits into positive- and negative-shifted sums $D^+$ and $D^-$, used identically in the forward and backward pass, preventing their roles from being isolated. We separate them into a family of operators that independently choose each denominator. The failure originates at initialization: every layer contains rows where $D^+$ is exactly zero, while near-dead rows produce gradient norms above $10^{12}$ regardless of the backward denominator. Only Softpick and a stop-gradient variant, which keeps $D^+ + D^-$ forward but backpropagates through $D^+$ alone, train from scratch. At 230M parameters, the stop-gradient operator matches Softpick on quantization, has fewer dead heads, and retrieves passkeys more reliably, trailing only on peak attention-weight kurtosis.

Comment: 23 pages, 4 Figures

arXiv abs page · PDF · same-day batch