PaperScope
LIVE · 2026-10-01 05:40 UTC

Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does

Jiaju Wu, Yi Hu, Muhan Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39892 v1
Category
Submitted
2026-09-30

Abstract

Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises a basic question: is each loop actually repeating the same computation, and if not, what routes the shared parameters to different operations? We study this question using graph walks as a test case. In the model's native trajectories, decoded predictions can advance by different numbers of graph steps or remain at a reached target, showing that recurrent progress need not follow a fixed one-loop-one-step pattern. We then show that a frozen loop can be steered toward different transitions by modifying its entering hidden state: a learned linear layer $J$ selects the desired transition without changing the shared Transformer layers. To test how this steering works, we use activation patching and find that attention patterns can recover its effects and switch the selected transition. Across five matched pairs of graph models, changing intermediate supervision during backbone training changes which transitions $J$ can induce. This suggests that $J$ selects computations learned by the backbone rather than creating new algorithms. Together, these results show that the hidden state can control shared computation, with attention routing as a causal pathway.

Comment: 36 pages. Code and reproduction materials: https://github.com/wjjpku/howloop

arXiv abs page · PDF · same-day batch