PaperScope
LIVE · 2026-09-29 05:40 UTC

Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33702 v1
Category
Submitted
2026-09-27

Abstract

Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.

Comment: 18 pages, 9 figures

arXiv abs page · PDF · same-day batch