PaperScope
LIVE · 2026-10-07 05:40 UTC

Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction

Lichen Zhu, Yueqian Lin, Yiheng Wang, Hai "Helen" Li, Yiran Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.08663 v1
Category
Submitted
2026-10-06

Abstract

Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the frames on which the estimator is above chance on average, so rejected frames reduce to the base exactly, while midrank normalisation lets an all-zero prior abstain at zero parameters. Gated fusion is significantly positive on film, sports and web video, whereas unconditional fusion is harmful on sports and null on web. Added to four supervised predictors, the NTIRE 2026 champion among them, FocusGate improves all sixteen model-domain cells in shuffled AUC, fifteen significantly, one domain pre-registered and scored once, while adding only 1% to the champion's latency. Alone, it surpasses TASED-Net and UNISAL in shuffled AUC on film with a 16-frame causal mean.

arXiv abs page · PDF · same-day batch