PaperScope
LIVE · 2026-09-10 05:40 UTC

Improving Cross-Lingual Token Representations by Adding a Pinch of SALT

Guillem Ramírez

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.09953 v1
Category
Submitted
2026-09-09

Abstract

Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to token-level tasks such as hallucination detection and sequence tagging, exposing a mismatch between training and usage. We propose SALT, a lightweight post-training method that improves token representations by injecting span-level supervision into existing sentence encoders. Across five multilingual token-level benchmarks, SALT achieves the best overall results on four of them, outperforming alternative fine-tuning strategies and competitive encoders. It also improves sentence-level performance on cross-lingual retrieval and classification tasks. These results demonstrate that span-level supervision is an effective signal for improving both token and sentence representations.

arXiv abs page · PDF · same-day batch