PaperScope
LIVE · 2026-09-22 05:40 UTC

Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy

Fengze Jia

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23048 v1
Category
Submitted
2026-09-19

Abstract

Compressed manipulation policies can pass offline evaluation while failing in closed-loop execution; this dissociation is established in prior work and is not our claim. We contribute a causal anatomy of one naturally occurring case. An 8-layer distillation of Octo-Base retains 86% of parameters, passes every offline check we applied (0.996 and 1.000 teacher-ratios on the family's own validation metrics), and collapses in closed loop: 0/72 vs. the teacher's 40/72 on a simulated WidowX pick-and-place task. The collapse is structured, not diffuse: early task stages degrade gradually (the student moves the object at 90% of the teacher's rate and grasps at 55%), while transport-to-target fails categorically, at 0% in every training variant. Paired action-trace forensics isolate the signature: a negative, late-heavy $z$ residual, roughly 10x its post-repair magnitude, and persistent across the base distillation and both continuation branches. Four standard therapies fail under matched controls: continued training and in-domain offline data leave success at zero, even though the latter measurably improves marginal action statistics; command-level compensation recovers nothing at any offset, although the same perturbations degrade healthy policies; clamping the symptom in the command channel preserves grasping, yet success stays at floor. A minimal-pair intervention that substitutes half of the training stream with deployment-distribution teacher rollouts, with every other setting held fixed, restores parity with the teacher (18/36 vs. 17/36 held-out), eliminates that signature, and recovers a teacher-like perturbation-response profile. We claim existence, not universality. Operationally, offline gates, including a family's own validation metrics, are insufficient acceptance tests for compressed policies; a few dozen closed-loop trials sufficed to find what they missed.

Comment: 8 pages, 1 figure, 5 tables. Accepted as a poster at the IROS 2026 Workshop on Building Scalable Infrastructure for Robot Learning: From Data Scaling to Real-World Deployment (ScaleInfra), Pittsburgh, PA, USA. Supplementary records: https://github.com/Xanadum/closed-loop-collapse

arXiv abs page · PDF · same-day batch