PaperScope
LIVE · 2026-10-06 05:40 UTC

Separating Decision Time from Decision Quality in the Real-Time Gap of Distilled Deciders: Evidence from a Game and a Conveyor Simulator

Chihoon Shin, Junyeong Lee, Kihyeok Jeong, Wonok Kwon

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04810 v1
Category
Submitted
2026-10-03

Abstract

Real-time agents are often evaluated with the world paused, or with decision time rounded to whole ticks (tick conversion). Tick conversion predicts almost no loss for a learned decider whose hit accuracy falls by 14.8 percentage points (pp) on the same games in asynchronous play. We instead split the decider's gap to a zero-latency teacher into time and quality components. In ViZDoom, we train a small decider by imitating a scripted teacher, run the game on wall-clock time, and add a control in which the teacher waits for a call to the decider's server before deciding. On a Windows host, three preregistered studies put the time component at 6.5-8.7 pp and the quality component at 6.3-8.6 pp. On a Linux host that skips 0.08-0.09% of ticks, the time component disappeared (-0.1 pp; paired 95% bootstrap interval [-0.4, 0.0]) while the quality component remained (5.6 pp [3.7, 7.5]). Retraining on teacher-labelled states from the decider's own play (DAgger) improved it on new games on both hosts (2.6 pp [0.4, 4.9] and 3.1 pp [1.1, 5.2]), a preregistered partial success. Imposing one delay schedule on every arm in live play left a quality component of 5.8 pp [4.0, 7.5] on 100 new games, and retraining cut the decider's disagreement with the teacher on its own states from 22.7% to 14.9%. In a conveyor simulator with a 400 ms deadline, delay past it erased the quality component. A latency-matched control whose delays match the decider's shows which remedy to try.

Comment: 80 pages including supplementary material. Submitted to Knowledge-Based Systems

arXiv abs page · PDF · same-day batch