PaperScope
LIVE · 2026-09-07 05:40 UTC

MoirfEolas and CríochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology

Jane Adkins, Abigail Walsh, Brian Davis, Elaine Uí Dhonnchadha

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.05022 v1
Category
Submitted
2026-09-04

Abstract

This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphological boundaries of the language. We present MoirfEolas, a dataset of over 35,000 Irish words mapped to their respective eclipses, prefixes and suffixes as well as an evaluation metric CríochScore, that evaluates the alignment of tokenizations with the morphological boundaries present in MoirfEolas. We evaluate common tokenization algorithms using CríochScore as well as intrinsic metrics present in the tokenization literature. We find that the Unigram Language Model aligns with Irish morphology more often than the other algorithms evaluated. We also find trade-offs between morphological-alignment of tokenization with both compression as well as vocabulary efficiency, providing practical insights for Irish natural language processing development. This dataset contributes towards combating the Irish language's low-resource status; moreover, the construction process reported in this paper can be emulated by other languages to create specialised morphological resources.

Comment: Accepted as a non-archival poster at the Second Tokenization Workshop (TokShop) at COLM 2026

arXiv abs page · PDF · same-day batch