PaperScope
LIVE · 2026-09-15 05:40 UTC

ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation

Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.13857 v1
Category
Submitted
2026-09-12

Abstract

Multimodal emotion recognition in conversation (ERC) requires adapting to the instance-dependent reliability of different evidence sources. Lexical content may be decisive, vocal expression may provide complementary cues, or accurate recognition may require cross-modal interaction; fixed fusion does not explicitly account for this variation. We propose ReH-FUSE, a reliability-aware framework with dialogue-aware text, audio, and cross-modal experts. Its decision-level router first models the relative preference between text and audio and then balances the resulting unimodal mixture against the cross-modal expert. This factorization separates unimodal competition from cross-modal selection. Across three independent runs on IEMOCAP, ReH-FUSE achieves 74.34% weighted F1 and 73.11% macro F1; on MELD, it achieves 68.03% weighted F1. Controlled ablations show that learned routing outperforms uniform expert averaging and benefits from cross-modal interaction.

Comment: 5 pages, 2 figures, 6 tables. Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch