Anonymous project page · Paper under double-blind review

Event-Localized Multi-Turn Reinforcement Learning
for Real-Time MOBA Commentary

A 9B vision-language model for real-time MOBA commentary on infinite video streams.

01Live demo

League of Legends

Given a 1 FPS video stream, MOBA-VL generates one commentary turn per second using only the frames it has seen so far.

02Live demo

Dota 2

The same streaming setup on professional Dota 2 matches.

03Live demo

Honor of Kings

The same streaming setup on professional Honor of Kings matches.

04Head-to-head

MOBA-VL vs. Proact-VL

MOBA-VL focuses on what a human caster would call out, such as how the fight unfolds, kills, dragons, and tower pushes. The prior method often falls back on generic filler right when a fight needs commentary.

Both models process video at 1 FPS, with their image-resolution settings (including min/max pixel limits) adjusted for 720p input.

05Paper

Abstract

Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released.