PaperScope
LIVE · 2026-09-11 05:40 UTC

ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations

Jiawen Wang, Kevin Yao, Khalid Jawed

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.10918 v1
Category
Submitted
2026-09-10

Abstract

Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-world scenes with unstructured obstacles remains a key generalization challenge. We present ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder. ObstaDiff extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. We evaluate ObstaDiff on 61 real-robot greenhouse trials per method (366 executions in total). ObstaDiff achieves 75.41% average task success and 8.20% average obstacle collision rate, outperforming representative imitation-learning baselines and improving generalization in cluttered agricultural scenes.

Comment: Accepted to the 10th Conference on Robot Learning (CoRL 2026), Austin, TX, USA. 16 pages, 5 figures

arXiv abs page · PDF · same-day batch