PaperScope
LIVE · 2026-09-29 05:40 UTC

From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features

Dewen Liu, Zixuan Li, Jonathan Pan, Zhao Wu, Zijun Yao, Juanzi Li, Xiaozhi Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35367 v1
Category
Submitted
2026-09-28

Abstract

Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.

Comment: 25 pages

arXiv abs page · PDF · same-day batch