PaperScope
LIVE · 2026-09-03 05:40 UTC

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context

Skanda Athreya, Yutong Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.01311 v1
Category
Submitted
2026-09-01

Abstract

We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the multiclass setting. This closes a gap left by prior work, whose multiclass result relied on a non-standard rounding-based approach rather than the typical argmax head used in practice.

arXiv abs page · PDF · same-day batch