PaperScope
LIVE · 2026-09-29 05:40 UTC

RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation

Jiawei Zhang, Xiangrong Zhang, Rui Song, Huanbin Zhou, Chengye Song, Hongzhou Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32416 v1
Category
Submitted
2026-09-26

Abstract

Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thus misassign training credit on states where the teacher fails. We propose RE-0, a recursively verified policy improvement framework: rather than assuming that the teacher globally outperforms the student, RE-0 requests local corrections from the teacher on the student's own failure histories and checks in the environment whether each correction is genuinely beneficial; verified corrections yield immediate improvement. Building on this, we propose RE-OPD, which turns verified interventions into supervision for on-policy distillation. Only counterfactually verified teacher interventions provide distribution-level supervision, weighted by their measured local benefit, and the improvement they induce is projected back into the standalone student, so both where supervision is applied and how much credit the teacher receives co-evolve with the student policy. We further prove that the student's per-round gain is lower-bounded by its verified intervention gain up to verification and projection error terms. Experiments across multiple Code-as-Policy embodied tasks show that RE-0 improves both teacher-assisted execution and the standalone student, and generalizes to novel robots and scenes.

Comment: 17 pages, 5 figures. Corresponding author: Hongzhou Wang (wanghongzhou@jlu.edu.cn)

arXiv abs page · PDF · same-day batch