PaperScope
LIVE · 2026-10-02 05:40 UTC

Do Multilingual Encoders Produce Language-Consistent Semantic IDs?

Abhinav Bohra, Anuj Bohra

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.01139 v1
Category
Submitted
2026-10-01

Abstract

Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.

Comment: 7 pages, 8 tables. Accepted as a short paper at WiNLP 2026, co-located with EMNLP 2026

arXiv abs page · PDF · same-day batch