PaperScope
LIVE · 2026-09-17 05:40 UTC

A Probe Shift Is Not a Fairness Fix: The Limits of Representation Steering in Speech Models

Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.18533 v1
Category
Submitted
2026-09-16

Abstract

Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretrained ASR encoders yield useful directions for reducing group word-error-rate (WER) gaps. Across Whisper-medium, HuBERT-large, and Wav2Vec2-large on Common Voice and the Speech Accent Archive, we probe every encoder layer for metadata-derived sex/gender, age, and native/accent labels; construct centroid and probe-derived directions; inject them at selected layers; and compare downstream probe trajectories with matched WER changes. Sex labels are highly decodable (best macro-F1 0.924--0.941), native/accent labels are also above chance (0.544--0.696), and age is weaker (0.354--0.397). Of 22 post-selected reruns, nine have 95% paired-bootstrap intervals entirely below zero, yet every absolute source-group WER reduction is below 0.7 percentage points. Conversely, a local target-class probe rate can rise from 8.09% to 99.87% while WER worsens. Linear readability is therefore neither evidence of causal use nor a reliable mitigation method. Our results motivate evaluating speech-bias interventions jointly at representation, propagation, and task levels.

Comment: Accepted at IMPACT-SPEECH 2026

arXiv abs page · PDF · same-day batch