PaperScope
LIVE · 2026-10-06 05:40 UTC

Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

Jie Wang, Shiwei Luo, Qi Zhang, Yuanbin Wu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05978 v1
Category
Submitted
2026-10-05

Abstract

Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to $25\%$ of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields $3.4\times$ more accepted tokens than in subword Transformers.

arXiv abs page · PDF · same-day batch