PaperScope
LIVE · 2026-10-01 05:40 UTC

Recovering Off-Policy Supervision for Speculative Decoding

Jungseob Lee, Chanjun Park, Sugyeong Eo, Hyeonseok Moon

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38795 v1
Category
Submitted
2026-09-30

Abstract

Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at https://github.com/js-lee-AI/ALR-IRA.

Comment: 22 pages, 4 figures, 17 tables

arXiv abs page · PDF · same-day batch