PaperScope
LIVE · 2026-09-29 05:40 UTC

Small transformers track Bayesian evidence for latent common causes via a context-invariant mechanism

Amir Mohammadpour, Michael Franke

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35161 v1
Category
Submitted
2026-09-28

Abstract

We present an in-depth investigation of how a form of Bayesian reasoning about common causes can emerge as a cross-contextual generalization in small, tractable transformers. Incrementing on recent work, our set-up (i) disentangles causal mechanisms in the model from the causal structure of the true data-generating process, (ii) orients more towards natural language prediction by considering inference of latent common causes, and (iii) considers whether and how Bayesian evidence accumulation for latent common causes can be implemented in representations and mechanisms that allow for cross-context generalization to novel test cases.

arXiv abs page · PDF · same-day batch