PaperScope
LIVE · 2026-09-15 05:40 UTC

Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs

JuHeon Ha, Byounghan Lee, Yunseo Choi, Kyung-Ah Sohn

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.15654 v1
Category
Submitted
2026-09-14

Abstract

Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimensions that need not correspond to independently controllable activation directions. Using the EPITOME framework, which decomposes supportive empathy into Emotional Reactions, Interpretations, and Explorations, we study three instruction-tuned LLMs and ask whether candidate directions derived from these labels produce distinguishable intervention effects or instead share structure, and how persona prompts interact with those directions. We find that contrastive activation addition yields a stable middle-layer intervention that consistently shifts the EPITOME proxy scores across models, moving empathy analysis beyond response-level scoring. However, the recovered directions are only partially separable: steering one direction induces off-target shifts, and hand-crafted prompting shifts the empathy profile rather than isolating a single dimension. Persona prompts substantially change EPITOME scores, but a paired activation-shift decomposition shows that the recovered subspace captures only approximately 3 percent of persona-induced squared activation-shift magnitude at layer 15. Under this EPITOME-based definition, expressed empathy is steerable but multi-axial, and controlling persona-conditioned empathy requires targeting structure beyond individual mechanism directions.

Comment: 18 pages, 6 figures. Accepted to the Main Conference of EMNLP 2026

arXiv abs page · PDF · same-day batch