PaperScope
LIVE · 2026-09-15 05:40 UTC

Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

Jaehun Shon, Jinha Choi, Jongwook Jeon, Jongmin Lee

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.15883 v1
Category
Submitted
2026-09-14

Abstract

Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies can naturally represent such multimodal behaviors, but learning an efficient one-step flow policy remains challenging: standard value guidance often leads to mode collapse or exploits overestimation bias in out-of-distribution regions. To address this, we introduce One-step Flow policy via Optimal Transport (OptiFlow), a framework for one-step flow policy learning as a structured sample-allocation problem. OptiFlow jointly trains a value-aware reference flow policy and an efficient one-step policy, coupling their action samples through state-wise entropic optimal transport. For each state, critic-estimated values define the priority of distillation target actions, while the action-distance cost ensures geometrically compatible pairings. By avoiding direct critic maximization, our transport-guided approach enables in-distribution exploitation by anchoring the one-step policy to high-value, dataset-supported modes without the risk of out-of-distribution divergence. Experimental results demonstrate that OptiFlow effectively captures optimal multimodal behaviors and achieves strong performance across diverse offline RL benchmarks. Our code is available at https://github.com/Yonsei-DILLab/OptiFlow.

Comment: Preprint, 38 pages

arXiv abs page · PDF · same-day batch