WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies
Lin Liu, Lu Zhang, Ziying Song, Wu Yang, Yuzheng Zhuang, Yunzhi Zhuge, Shuai Tao, Wulong Liu, Huchuan Lu
Abstract
Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed interactions. We propose \textbf{WorldGuide}, a framework that learns these distinctions in latent space and uses them to guide policy training. WorldGuide combines predictive pretraining on successful and failed trajectories with contrastive learning on matched success--failure pairs. The learned predictor then provides a differentiable reward to guide joint optimization of the policy and visual encoder. The predictor is discarded after training, so deployment requires no additional world-model inference. Extensive experiments show that WorldGuide substantially improves VLA reliability and achieves state of the art performance on LIBERO 100 and SimplerEnv, reaching \textbf{96.8\%} and \textbf{72.0\%}, respectively. Code will be publicly available.