PaperScope
LIVE · 2026-09-22 05:40 UTC

Vision-Wireless Fusion for Multi-User Localization: A Cross-Modal Transformer Approach

Can Zheng, Jiguang He, Guofa Cai, Henk Wymeersch, Merouane Debbah

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23372 v1
Category
Submitted
2026-09-20

Abstract

Accurate multi-user localization is challenging in complex urban environments, where wireless measurements can become ambiguous under noise, blockage, and multipath, while visual observations provide complementary spatial context. This paper presents a vision-wireless fusion framework for multi-user localization using pilot-indexed channel state information (CSI). Orthogonal pilot indices preserve the identities of the communicating UEs in the CSI token sequence and localization outputs. The model encodes each pilot-indexed CSI observation as a query token and uses cross-attention to retrieve user-specific information from spatial visual memory. Self-attention among CSI tokens further captures inter-user interactions, while the resulting multimodal representations are used for user-wise localization. Experiments on different datasets show consistent improvements over model-based, CSI-only, and multimodal-fusion baselines. Further experiments evaluate the model under different wireless and visual conditions.

Comment: 12 pages, 8 figures, 7 tables

arXiv abs page · PDF · same-day batch