PaperScope
LIVE · 2026-10-01 05:40 UTC

Completion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding

Yifan Zhang, Qifan Zhang, Liang Zheng

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39868 v1
Category
Submitted
2026-09-30

Abstract

Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with simulator interaction, but the inexpensive online simulator can differ from the target in transition and event-duration dynamics. We study this cross-fidelity problem for multi-line bus holding and address a failure mode in which lower generalized passenger time coexists with incomplete passenger journeys.

arXiv abs page · PDF · same-day batch