PaperScope
LIVE · 2026-09-15 05:40 UTC

Towards Identifying the Dataset Biases Causing Phantom Transfer

Jonas Jürß, Pietro Liò

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.14449 v1
Category
Submitted
2026-09-13

Abstract

Recent work has shown that a teacher model can transfer a bias to a student through a dataset from which every explicit reference to that bias has been filtered out, and that no data-level defense reliably removes or detects it even when knowing what bias to look for. Aiming to shed light on the hidden traces of these biases, we show that a simple signature based on Sentence BERT embeddings can identify the topic of such a bias with a Matthews correlation coefficient of 0.83 if the teacher model used by the attacker is known and 0.46 if it is not. Additionally, we observe that different teacher models appear to express the same bias through different vocabulary.

arXiv abs page · PDF · same-day batch