PaperScope
LIVE · 2026-09-29 05:40 UTC

GSM: Efficient Language Modeling with Shared Global State

Yunao Zheng, Bin Wen, Xiaojie Wang, Kaiyu Jiang, Xuanyu Zheng, Changyi Liu, Hongyi Fu, Jianxiong Wang, Tianke Zhang, Haonan Fan, Yingxin Li, Jiankang Chen, Xu Wang, Tingting Gao, Han Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33465 v1
Category
Submitted
2026-09-27

Abstract

Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.

arXiv abs page · PDF · same-day batch