PaperScope
LIVE · 2026-10-08 05:40 UTC

Document-Level Text Simplification in Estonian Using Large Language Models

Meeri-Ly Muru, Eduard Barbu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10378 v1
Category
Submitted
2026-10-07

Abstract

Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.

Comment: 12 pages, 2 figures, 2 tables. Published at LREC 2026

Journal: Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pp. 7225-7235, 2026

arXiv abs page · PDF · same-day batch