PaperScope
LIVE · 2026-10-09 05:40 UTC

WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models

Kai Ding, Yang He, Ruijie Quan, Yi Yang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.11401 v1
Category
Submitted
2026-10-08

Abstract

World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.

Comment: 19 pages, 5 figures, 7 tables. Project page: https://dingkai0302.github.io/wam-cache/

arXiv abs page · PDF · same-day batch