PaperScope
LIVE · 2026-09-03 05:40 UTC

Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

Leandra Fichtel, Janek Prange, Henning Wachsmuth

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29215 v2
Category
Submitted
2026-08-29

Abstract

To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. Prompting alone has been shown to be insufficient for creating such explanations, and other computational methods are missing so far. Therefore, this paper investigates whether LLMs can be steered to generate explanations that are tailored to a specific group of people. To this end, we propose an approach that first identifies group-specific attributes in terms of explanatory style and knowledge of a specific target group. Building on activation engineering, it then computes attribute-based steering vectors and adds them to the internal activations of an LLM during inference to enable a fine-grained steering. In our experiments, we assess the steering effectiveness in terms of specificity and factuality of the generated explanations. Additionally, we evaluate the explanations in a study with human experts from different target groups. Compared to prompting and state-of-the-art steering baselines, our approach tailors the explanations significantly better to the target group while maintaining the best specificity-factuality balance.

Comment: Accepted to EMNLP 2026 Main

arXiv abs page · PDF · same-day batch