PaperScope
LIVE · 2026-09-17 05:40 UTC

LocQE: Principled Domain Adaptation for Localisation Quality Estimation by Leveraging Post-Edits

Kathy Hämmerl, Gabriel Bretschner, Joern Wuebker

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.18720 v1
Category
Submitted
2026-09-16

Abstract

Learned quality estimation (QE) models such as COMETKiwi are widespread and work well for general machine translation evaluation. However, they are known to struggle on unseen domains, limiting their performance in a real-world localisation context. We show that they are insensitive to some important factors in localisation, such as whether numbers are translated accurately, or even whether the correct number of spaces and punctuation are preserved in a translation. Further, a key capability for optimisation of machine translation is the ability of QE models to accurately rank different translations of a single segment, which suffers significantly from the domain transfer. In the absence of large-scale direct assessment data, we propose principled fine-tuning approaches to reduce the domain gap with even small amounts of post-editing data. Using a multi-task fine-tuning approach and a simple tokeniser intervention, we create a QE model which proves markedly better at distinguishing preferred post-edits from rejected initial translations in a localisation context. We show that preferences and artificial continuous scores stabilise each other, and argue that to calibrate metrics both in terms of their absolute scores and comparisons between translation of the same source, both types of signal are needed.

arXiv abs page · PDF · same-day batch