PaperScope
LIVE · 2026-10-08 05:40 UTC

Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

Xiaoran Liu, Ziwei He, Xipeng Qiu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10114 v1
Category
Submitted
2026-10-07

Abstract

The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16$\times$ training-free length extrapolation while maintaining 100\% accuracy on NIAH-SK1 in 64k context length.

Comment: 60 pages, 36 figures, 25 tables, under review

arXiv abs page · PDF · same-day batch