PaperScope
LIVE · 2026-10-01 05:40 UTC

Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering

Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie Song

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39701 v1
Category
Submitted
2026-09-30

Abstract

Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.

arXiv abs page · PDF · same-day batch