PaperScope
LIVE · 2026-09-03 05:40 UTC

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

Jungseob Lee, Jaehyung Seo, Heuiseok Lim

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.28930 v1
Category
Submitted
2026-08-28

Abstract

Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full-dimensional classifiers, so apparent architectural complexity largely reflects high-dimensional covariance estimation difficulty rather than exploitable non-linearity. A simple L2-regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi-layer aggregation exceeds CLAP cross-layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle-layer performance without oracle access. Our claims characterize the geometry within the controlled paired-example paradigm. Our code is available at https://github.com/js-lee-AI/LayerMix.

Comment: 19 pages, 7 figures, 20 tables. Accepted to EMNLP 2026 (Main Conference). Code: https://github.com/js-lee-AI/LayerMix

arXiv abs page · PDF · same-day batch