PaperScope
LIVE · 2026-09-28 05:40 UTC

Low-Bit Recurrent States in Hybrid Language Models

Hongren Chen, Jiayang He

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.30950 v1
Category
Submitted
2026-09-25

Abstract

Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative log-likelihood by factors of 3.3--27.9 relative to the best of seven baselines across three hybrid models; metadata costs vary. At six bits, negative log-likelihood differs from the FP32-state baseline by less than 0.005 nats. Ablations separate gains from variable bit widths, decay weighting, and range normalization. With less frequent write-backs, gains diminish and depend on the model and budget.

Comment: 9 pages, 2 figures, 3 tables. Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch