PaperScope
LIVE · 2026-09-03 05:40 UTC

MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?

Uicheol Jung, Juyoung Hong, Geuntaek Lim, Yukyung Choi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.02565 v1
Category
Submitted
2026-09-02

Abstract

Text-video retrieval requires representations that can distinguish videos with similar scenes, actions, and temporal patterns. Recent multimodal large language models have been adapted as embedding models, but they often represent each input using a single token from the final layer. This can compress diverse video-text cues into a single vector and limit fine-grained retrieval. To address this limitation, we propose MARS, a multi-layer and multi-slot embedding framework for text-video retrieval. MARS constructs multiple adaptive representation slots by combining hidden states from different decoder layers, compares corresponding text and video slots, and aggregates their similarities for retrieval. To better handle confusing candidates, we further introduce a hard-negative-aware slot specialization objective that encourages the slots to capture discriminative matching cues. Experiments on four text-video retrieval benchmarks show that MARS achieves state-of-the-art results in both direct similarity-based retrieval and reranking settings. Ablation studies and analyses demonstrate that multi-layer fusion, multiple slots, and hard-negative-aware slot specialization provide complementary gains. Code is available at https://github.com/sejong-rcv/MARS.

Comment: 16 pages, 6 figures. Accepted to the Main Conference of EMNLP 2026

arXiv abs page · PDF · same-day batch