PaperScope
LIVE · 2026-09-17 05:40 UTC

ReMova: Fine-tuning LLMs for English to Belarusian translation

Mikita Pilinka, Aliaksandr Kliujeŭ, David Samuel, Yves Scherrer

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.16427 v1
Category
Submitted
2026-09-14

Abstract

This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation. Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the issue of the two orthographies of the Belarusian language, noise in the training data, interference from other languages and other misspelling issues common in Belarusian on the internet. A matched ablation on unfiltered training data shows substantial benefits from filtering for all fine-tuned models, with the LLM-based models gaining roughly twice as much from filtering as the dedicated encoder-decoder MT system, supporting the view that for Belarusian MT one of the primary bottlenecks is data quality.

Comment: WMT26 submission

arXiv abs page · PDF · same-day batch