PaperScope
LIVE · 2026-10-05 05:40 UTC

Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.03333 v1
Category
Submitted
2026-10-02

Abstract

Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: https://vista-paper.github.io/

Comment: 21 pages, 6 figures. Accepted to the 10th Conference on Robot Learning (CoRL 2026)

arXiv abs page · PDF · same-day batch