PaperScope
LIVE · 2026-10-01 05:40 UTC

MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary

Shengyun Zhong, Xinkang Zhao, Ziyuan Chu, Linchao Zhu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38428 v1
Category
Submitted
2026-09-29

Abstract

Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at https://moba-vl.github.io.

Comment: 30 pages, 12 figures

arXiv abs page · PDF · same-day batch