PaperScope
LIVE · 2026-09-07 05:40 UTC

Interpretability for Turing Machines

Billy Snikkers, Rumi Salazar, Daniel Murfet, Will Troiani

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.04661 v1
Submitted
2026-09-04

Abstract

We show that susceptibilities, an interpretability technique developed for neural networks, can identify the presence of algorithmic structure in Turing machines by probing the local loss landscape of a learning problem for noisy Turing machines introduced by Murfet and Troiani (arXiv:2504.08075). We prove that symmetries and path separation in the algorithm implemented by a Turing machine induce permutation symmetries and low-rank blocks in its susceptibility matrix. We study this empirically on a set of deterministic finite automata (DFAs) and demonstrate that algorithmic features can be recovered by principal component analysis and clustering methods in susceptibility space.

Comment: 75 pages, 31 figures, 3 tables. Interactive companion: https://tminterp.timaeus.co. Code and data: doi:10.5281/zenodo.22205895

arXiv abs page · PDF · same-day batch