PaperScope
LIVE · 2026-09-29 05:40 UTC

Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs

Fansheng Zhang, Shengran Guo, Zexiao Wang, Liang Yuan, Jiyuan Chen, Ruikun Luo

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33298 v1
Category
Submitted
2026-09-27

Abstract

In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.

Comment: 32 pages, 7 figures. Code: https://github.com/abovefiramament/DHSA

arXiv abs page · PDF · same-day batch