PaperScope
LIVE · 2026-09-09 05:40 UTC

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

Gonçalo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos, Duarte Miguel Alves, Afonso Simplício, Diogo Tavares, David Semedo, Daniel Gomes, João Magalhães

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07699 v1
Category
Submitted
2026-09-07

Abstract

Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a clean, representative corpus optimized for LLM pre-training.

Comment: 16 pages, 9 figures, EMNLP 2026 Main

arXiv abs page · PDF · same-day batch