PaperScope
LIVE · 2026-09-22 05:40 UTC

Enhancing speech representation learning with cross-modal knowledge transfer with HGNN under low resource settings: the case study of Yemba

Yannick Yomie Nzeuhang, Paulin Melatagia Yonta, Marie Tahon

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23194 v1
Category
Submitted
2026-09-19

Abstract

Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiveness of traditional and self-supervised methods. As a promising alternative, in this work, we propose to enhance acoustic representation trough a cross-modal transfer knowledge approach, based on heterogeneous graph neural networks (HGNNs), where acoustic and linguistic entities are modeled as distinct node types within a unified graph. Through message-passing mechanisms, linguistic nodes explicitly transfer knowledge to acoustic nodes, enabling structured and interpretable cross-modal information flow. To highlight this knowledge transfer and its benefits, we measured standard clustering metrics as an intrinsic evaluation of acoustic representation, and to emphasize applicability, we performed isolated-word recognition tasks using an English benchmark and a Cameroonian language dataset in low resources settings . Results demonstrate that acoustic representations consistently benefit from linguistic knowledge propagated through the graph. To our knowledge, this is the first demonstration of explicit cross-modal knowledge transfer for acoustic representation learning using HGNNs, highlighting a promising direction for speech representation in low-resource settings.

arXiv abs page · PDF · same-day batch