PaperScope
LIVE · 2026-09-03 05:40 UTC

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.27510 v1
Category
Submitted
2026-08-27

Abstract

Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept-Targeted Attribution (CTA), which instead trains attribution graphs with respect to a linear probe direction. CTA therefore yields probe-specific circuits that explain why an internal concept representation arises in a prompt, independently of whether it is expressed in the generated token. Using Cross-Layer Transcoders, we show that these probe-targeted graphs contain predictive structure: graph-level features predict probe accuracy across four widely studied concept categories ($ρ= 0.91$, $R^2 = 0.84$), while local features identify the sparse components driving per-prompt classification. This connects probe performance to interpretable circuit structure, allowing us to ask not only whether a probe works, but which internal computations make it work. Causal ablations further show that probe-targeted and logit-targeted graphs capture functionally distinct mechanisms. Removing probe-relevant features reduces internal concept scores while largely preserving generated tokens, whereas removing logit-relevant features changes the generated token in 92% to 100% of cases with near-zero effect on probe scores. CTA provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones. Our code is available at https://github.com/vedantpalit/concept-targeted-attribution

Comment: Accepted to EMNLP 2026 (Main Conference). 29 pages, 18 figures, 9 tables

arXiv abs page · PDF · same-day batch