PaperScope
LIVE · 2026-09-21 05:40 UTC

World Modeling in Transformers

Pierre Beckmann, Matthieu Queloz, Andre Freitas

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.21748 v1
Category
Submitted
2026-09-18

Abstract

Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position, and uses a goal compass to navigate. We trace its failures to interference between superposed intersection features, which disrupts localization within the internal map. Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors. Finally, we propose mechanistic indicators that we use to compare models and show that world-modeling capacities emerge at different stages of training. Our findings motivate a shift from asking whether a model has a world model to mechanistically studying its world modeling: the interacting capacities through which it represents its environment and uses those representations to guide behavior.

arXiv abs page · PDF · same-day batch