PaperScope
LIVE · 2026-09-22 05:40 UTC

SPACE: Semantic Projection and Alignment of CLIP Embeddings for Domain Adaptation

João Renato Ribeiro Manesco, Danilo Samuel Jodas, Douglas Rodrigues, Leandro Aparecido Passos, João Paulo Papa

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23248 v1
Category
Submitted
2026-09-19

Abstract

A fundamental challenge in deploying vision models is domain shift, which arises when training and test data follow different distributions, leading to degraded performance. This challenge is amplified when the same semantic concept appears under distinct visual forms, such as photographs and sketches, where visual similarity is weak despite semantic correspondence. Existing unsupervised domain-adaptation methods aim to align distributions across domains but often ignore semantic relationships among samples of the same class. To address this issue, this paper introduces SPACE, a method that exploits the semantic structure of CLIP's vision-language space for domain adaptation. The key idea is to use text descriptions as semantic anchors by applying Singular Value Decomposition to CLIP embeddings of class descriptions, yielding an orthogonal basis that captures semantic relationships among categories. Visual features from both domains are projected into this semantic subspace, aligning images based on meaning rather than appearance.

Comment: Accepted for publication at the 2026 IEEE International Conference on Image Processing (ICIP), Tampere, Finland. 7 pages, 2 figures, 3 tables

Journal: 2026 IEEE International Conference on Image Processing (ICIP), Tampere, Finland, 2026

arXiv abs page · PDF · same-day batch