PaperScope
LIVE · 2026-09-29 05:40 UTC

Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers

Swagatam Mukhopadhyay, Vishal Vivek Saley, Vraj Parikh, Mausam

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35630 v1
Category
Submitted
2026-09-28

Abstract

Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks corrective feedback while allowing continued accumulation. We find that both attention and feed-forward layers have this read-blindness, and contribute to the emergence and persistence of MAs. To validate prior work that hypothesized that FFN's amplification abilities is the primary reason for MAs (Sun et al., 2026), we analyze the model checkpoints during learning. Contrary to our expectation, read-blindness emerges before FFN amplification, suggesting that it acts upstream in the MA mechanism. We further contribute gradient analysis to link this behavior to surprising asymmetries in the loss landscape, concluding that the model actively maintains this read-blindness. Finally, we find that removing read-blocking at different locations induces compensatory shifts elsewhere, but MAs still persist.

arXiv abs page · PDF · same-day batch