PaperScope
LIVE · 2026-10-01 05:40 UTC

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

Sheng Zhao, Weikai Lin, Yuhao Zhu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38519 v1
Category
Submitted
2026-09-29

Abstract

Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.

Comment: Accepted at NeurIPS 2026. 25 pages

arXiv abs page · PDF · same-day batch