PaperScope
LIVE · 2026-10-01 05:40 UTC

Why Do Conventional World Models Fail to Learn Cellular Automata?

Shaoyang Guo, Ziming Liu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39604 v1
Category
Submitted
2026-09-30

Abstract

Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in many cases. Conventional architectures predict most pixels correctly yet rarely complete a rollout: a CNN predicts 96.3% of cells but completes 18.9% of rollouts; a joint diffusion model completes none. We trace the gap to three failure modes of these world models - namely, they fail to exactly capture spatial locality, temporal locality or temporal stability. Simple changes repair each: (1) for spatial locality, two-dimensional rotary positions lift a transformer from 39.1% to 100% on the Game of Life; (2) for temporal locality, handing each token its cell's previous-frame neighbourhood lifts the same transformer from 25.8% to 99.9% on unseen rules; (3) for temporal stability, causal freezing lifts the same diffusion weights from 42.2% to 99.9%. None of the three changes touches the architectural backbone; each only modifies the information flow within it. We also compare joint and ordered sampling on billiards and, in an exploratory study, on a simulated Burgers equation.

Comment: 35 pages, 18 figures. Code and reproduction materials: https://github.com/guoshaoyang-pku/momentum-induction

arXiv abs page · PDF · same-day batch