PaperScope
LIVE · 2026-09-17 05:40 UTC

MEgoVista: Multi-view Ego-aware Motion Estimation for Metric 4D Hands and Head in the Wild

Jiangong Xiao, Zhihao Zhang, Yifei Dong, Chao Ma, Zhouyi Jin, Zhiwen Hou, Li Liu, Weihuang Chen, Hongbin Sun, Maoqing Yao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.16684 v1
Category
Submitted
2026-09-15

Abstract

Learning manipulation from human video requires high-fidelity hand-motion reconstruction in metric units. Today's metric hand labels come from studio rigs and instrumented headsets, and both are confined in the same two ways: neither leaves a prepared setting, and neither is checked against an independent reference. Unconstrained head-worn recording promises the opposite trade-off, scaling with the number of people wearing a device. We therefore introduce MEgoVista, an offline pipeline that turns a single unprepared MEgo View recording into metric two-hand and head motion in one gravity-aligned world frame. Three properties set it apart from existing egocentric reconstruction systems: first, it reconstructs in settings studio volumes and tabletop rigs cannot reach, settling hand ownership at detection so bystander hands stay out of the wearer's trajectory; second, it takes its metric gauge from calibrated stereo rather than a monocular prior, installing scale at initialisation so policies receive physical units, not arbitrary coordinates; third, both outputs are scored inside a motion-capture volume against independent Chingmu optical capture, under a protocol that audits its own reference and charges what a method declines to predict. MEgoVista is offered as a measured route from egocentric video to metric hand supervision, one that widens where such labels can be gathered.

Comment: 13 pages, 3 figures, 3 tables

arXiv abs page · PDF · same-day batch