PaperScope
LIVE · 2026-10-01 05:40 UTC

TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection

Bo-Yuan Cheng, Kuan-Yu Chen, Po-Han Huang, Jeng-Lin Li, Jian-Jiun Ding

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39051 v1
Category
Submitted
2026-09-30

Abstract

Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal missingness, where frames are missing independently in each modality, and stream-level missingness, where one modality is unavailable throughout a video. Moreover, we find that the mean squared error (MSE) loss is misaligned with both the evaluation metrics and the peak-driven nature of highlights. Therefore, we propose Temporal-Stream Modality Dropout (TSMD), which combines structured missingness simulation with a joint objective comprising pointwise MSE, per-video Pearson correlation, and peak-oriented RankNet loss terms. TSMD has three variants: temporal, stream-level, and mixed dropout. On the MoSu and Mr. HiSum datasets, TSMD-Temporal improves mAP@15 by 7.06 and 3.41 points over TripleSumm under 50% independent temporal removal, whereas TSMD-Stream performs the best under complete-stream removal. TSMD-Mix retains most of these complementary benefits and ranks the best or the second-best across the evaluated temporal and stream-level conditions.

Comment: 5 pages, 3 figures, 3 tables

arXiv abs page · PDF · same-day batch