PaperScope
LIVE · 2026-10-02 05:40 UTC

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

Kailen Hargenrader, Edoardo Calvello, Bohan Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.00564 v1
Category
Submitted
2026-09-30

Abstract

Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.

Comment: 32 pages, 17 figures, accepted to NeurIPS 2026 Workshop on AI for Stochastic Dynamics

arXiv abs page · PDF · same-day batch