PaperScope
LIVE · 2026-09-09 05:40 UTC

TurEngMix: A Text Corpus and Benchmark for Turkish-English Code-Mixed Language Identification and Named Entity Recognition

Ilayda Dogan, Phuong-Anh Nguyen-Le, Julia Mendelsohn

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06963 v1
Category
Submitted
2026-09-07

Abstract

Natural language processing systems underperform on code-mixed text, particularly for low-resource language pairs. Turkish-English poses a further challenge: it lets English stems combine with Turkish suffixes to form single mixed-language tokens. We introduce TurEngMix, a corpus of 5.5K noisy, naturally occurring social media posts (486,974 tokens) rich in Turkish-English code-mixing. From this corpus, we construct a new Turkish-English benchmark for code-mixed language identification (LID) and named entity recognition (NER), comprising 15K expert-annotated tokens. Evaluating both decoder LLM and fine-tuned encoder baselines, we find that monolingual Turkish and English tokens are labeled reliably, but all models have high error rates on mixed-language tokens for both LID and NER. For morphologically integrated tokens, NER error rates were 5.2x and 6.3x higher for GPT-4o and Qwen, respectively. This highlights how morphological integration remains a challenge. We release the corpus, annotations, and code to support future computational and sociolinguistic research on Turkish-English code-mixing.

Comment: Accepted to W-NUT 2026 (11th Workshop on Natural User-generated Text), co-located with EMNLP 2026. 15 pages, 4 figures, 12 tables

arXiv abs page · PDF · same-day batch