PaperScope
LIVE · 2026-09-24 05:40 UTC

ThaiTrees: Thai Syntactic Dependency Trees Across Domains

Attapol T. Rutherford, Papatchol Thientong

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.27558 v1
Category
Submitted
2026-09-23

Abstract

Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research. We present ThaiTrees, a 342M-token corpus drawn from news, Wikipedia, spoken transcripts, and social media. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions. We release a frequency lexicon and CoNLL-U parses in machine-readable formats suitable for both AI-assisted and conventional programmatic analysis.

arXiv abs page · PDF · same-day batch