PaperScope
LIVE · 2026-09-17 05:40 UTC

ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

Zahra Bokaei, Walid Magdy, Bonnie Webber

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.16393 v1
Category
Submitted
2026-09-14

Abstract

We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identification across seven structured target categories. ParsHate also distinguishes explicit and implicit hate, marks explicit and implicit targets, and provides span-level rationales. Data collection combines random and score-stratified temporal sampling to reduce keyword-driven bias while preserving natural label distributions. Applying SOTA models for Persian hate-speech detection on ParsHate shows moderate performance (79% F1), especially with samples from earlier years, and low performance with target identification (25.5% macro-F1). This emphasizes the diverse sampling of hate speech in ParsHate and its challenging nature that requires more advanced methods for better performance. Dataset is made publicly available.

Comment: Accepted to EMNLP 2026 (Main Conference)

arXiv abs page · PDF · same-day batch