PaperScope
LIVE · 2026-09-09 05:40 UTC

The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs

Ziyue Feng, Hongbo Fang, James A. Evans

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07117 v1
Category
Submitted
2026-09-07

Abstract

Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language model says, but it is unclear which layer they reach. Do they reconfigure internal structure, or only modulate the output channel? We use persona conditioning as a controlled probe, measuring its effects along a depth axis from self-report, through open-ended generation, to word-level parametric association, across three instruction-tuned models. We find a graded dissociation. Personas are legible but not structural: models follow single-trait instructions yet fail to reproduce human inter-trait covariance. The dissociation deepens with depth: personas hold or amplify closed-form QA bias, shift absolute tone while leaving between-group disparity unchanged, and barely perturb an already saturated associative baseline. Prompt-based steering thus operates in the output channel and has a structural reach limit that surface manipulability can mask.

arXiv abs page · PDF · same-day batch