PaperScope
LIVE · 2026-09-30 05:40 UTC

Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features

Dae Ung Jo, Jongin Lim, YoungJoon Yoo, Daeho Um

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.37243 v1
Category
Submitted
2026-09-29

Abstract

Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.

arXiv abs page · PDF · same-day batch