PaperScope
LIVE · 2026-10-06 05:40 UTC

From Abusive Language Classification to Sequence Labeling Identification

Nicolas Zampieri, Ignacio Lopez, Manon Girard, Jeremy Auguste

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06287 v1
Category
Submitted
2026-10-05

Abstract

Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is targeted. We define Abusive Language Identification (ALI) as a sequence-labeling task that jointly extracts AL spans and target mentions, and assess whether this approach can be used for text moderation. On a pilot corpus drawn from a production moderation pipeline, we compare ALI with ALC on cross-domain generalization and implicit abuse, and we also evaluate AL and target span detection. ALI remains competitive with ALC while providing localized outputs for moderators, with a modest and configuration-sensitive advantage on implicit abuse. Exact AL boundaries and target spans remain difficult to recover. We complement this comparison with a qualitative analysis and discuss perspectives on complete target--span linking and on structured benchmarks for ALI.

arXiv abs page · PDF · same-day batch