PaperScope
LIVE · 2026-09-03 05:40 UTC

CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space

Paul Schneider, Nazim Haouchine

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.00272 v1
Category
Submitted
2026-08-31

Abstract

Most advances in keypoint descriptions address monomodal settings, where image variations arise from viewpoint, illumination, or contrast changes. Multimodal scenarios involve images produced by fundamentally different sensing processes, such as multispectral imaging, RGB-depth, satellite imagery, or medical imaging, causing the same structures to appear differently. A common solution to cross-modal description is to train descriptors for each modality pair, which requires retraining whenever the modalities change, or to train large models, which incur a significant increase in runtime. Instead, we propose CrossFeat, a framework that enables an existing monomodal descriptor to operate across modalities. Our method learns a crossing function in descriptor space that maps features from one modality to a representation compatible with another. To preserve the structural information captured by the original descriptor, CrossFeat introduces a geometry-appearance disentanglement such that only appearance is altered while the geometric properties are preserved. Experiments across multiple domains and datasets demonstrate improved performance in multimodal matching.

Comment: ECCV 2026

arXiv abs page · PDF · same-day batch