PaperScope
LIVE · 2026-09-03 05:40 UTC

TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection

Qianqian Chen, Hyun Bin Kim, Denzel Elden Wijaya, Yang Yi, Bo Liu, Yangkai Ding

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29577 v1
Category
Submitted
2026-08-30

Abstract

Traditional video highlight detection relies on a narrow, event-centric definition of saliency, which often fails to generalize to unconstrained personal videos where highlights are heterogeneous and perspective-dependent. To address this, we introduce TRINITY, a multi-perspective benchmark that decomposes highlight saliency into three complementary dimensions, Event, Emotion, and Nature, within a unified temporal framework. Leveraging this multi-faceted view, we propose a shared-backbone multi-branch architecture designed for parallel multi-perspective prediction via view-specific experts. Comprehensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines, achieving gains of +7.15/+3.62 mAP (rho=15%/50%) on Mr. HiSum and +10.82 mAP on YouTube Highlights. These results validate that multi-perspective modeling provides a more robust and comprehensive formulation of video saliency, especially for complex real-world scenarios. The benchmark and relevant codes will be released upon acceptance. The benchmark is available at https://huggingface.co/datasets/vanilladucky/TRINITY and the code is available at https://github.com/vanilladucky/TRINITY.

Comment: 32 pages, 9 figures. Accepted to ECCV 2026

arXiv abs page · PDF · same-day batch