PaperScope
LIVE · 2026-09-30 05:40 UTC

How Language Models Differ in Redistributing Attention-Head Activity Under Serial Demand

Johnny Jingze Li, Abdulla Kuleib, Kalyan Basu, Gabriel A. Silva

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36221 v1
Category
Submitted
2026-09-28

Abstract

The way a model distributes activity over each layer's attention heads offers a coarse view of how it routes information through depth; how this changes with the task is part of what a mechanistic account must explain. Holding prompt length fixed, we vary how many serial steps a task demands and measure, in every layer of 17 open-weight models, whether activity concentrates on a few heads or spreads across many as demand rises. Both occur: in most models, layers just before mid-depth concentrate activity and later layers spread it. Models differ in where and how strongly this happens. The Qwen2.5 base models from 0.5B to 7B, for example, spread less than the average model in every task and concentrate activity in parts of their second half, where Llama models from 1B to 8B and OLMo-2 spread; the contrast largely holds between Llama-3.1-70B and Qwen2.5-72B, which have the same number of layers and heads. These differences are reproducible, and post-trained models keep much of their base model's pattern. An ablation study suggests that, within a task, models whose activity is more concentrated on their top heads also depend more on those heads for the answer. Concentration and spreading across layers thus offer a new way to compare models, by how they route information through depth. Code is available at https://github.com/johnnyjli/serial-demand-heads.

arXiv abs page · PDF · same-day batch