PaperScope
LIVE · 2026-09-15 05:40 UTC

Calibrating Interpretability Instruments Before Trusting Their Verdicts

Orion Reblitz-Richardson

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.14754 v1
Category
Submitted
2026-09-13

Abstract

Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an interchange patch among others. These measurements fail in specific, diagnosable ways that return a plausible number instead of an error, so a broken instrument can easily read as a finding. A covariance-matched null can saturate until every direction looks typical, a per-head attribution can overshoot the true residual write threefold on reordered-normalization architectures, an interchange patch can go sign-chaotic because its outcome is pinned at a ceiling, or a read-from verdict can be an artifact of measuring past the layer where the model already decided. This note documents six such failures from a causal interpretability program on refusal and moral representation, spanning several papers and a four-model open-weight panel; each mode is established on one or two of the four. For each we give the tell that catches it and a protocol keyed to a detectable trigger (reordered normalization, massive activations, a low-dimensional decision channel), so we and readers can check whether a given setup is exposed. The discipline reduces to four moves: calibrate against a positive-control ladder, certify with an orthogonal cell, compute power before spending compute, and state every read-from verdict at a depth referenced to the model's commitment. The evidence is four architectures across three families within a single program; external replication across programs is future work.

Comment: 16 pages, 3 figures, 1 table. Companion to "Refusal Reads Only a Slice of What the Model Knows" (submitted concurrently). Per-unit arrays: https://doi.org/10.5281/zenodo.22731361. Code: https://github.com/deepsteer/deepsteer

arXiv abs page · PDF · same-day batch