PaperScope
LIVE · 2026-09-29 05:40 UTC

A Comparative Analysis of Attention versus State-Space Models for In-Context Learning

Enes Arda, Semih Cayci, Atilla Eryilmaz

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32341 v1
Category
Submitted
2026-09-26

Abstract

Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.

arXiv abs page · PDF · same-day batch