PaperScope
LIVE · 2026-09-28 05:40 UTC

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.30724 v1
Category
Submitted
2026-09-25

Abstract

Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.

Comment: 5 pages, 3 tables, 1 figure. Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch