PaperScope
LIVE · 2026-09-03 05:40 UTC

GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning

Outongyi Lv, Yuanwei Zhang, Xiaoqun Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.30632 v1
Category
Submitted
2026-08-31

Abstract

Reinforcement learning (RL), particularly RL with Verifiable Rewards (RLVR), has recently emerged as a central paradigm for enhancing large language models' (LLMs) reasoning abilities, demonstrating remarkable effectiveness across reasoning tasks. Recent studies suggest that high-entropy tokens play an exceptionally important role in model training, since training with only the highest 20% entropy tokens yields significant performance gains. However, why such high-entropy tokens are beneficial remains insufficiently understood. In this work, we find that although high-entropy tokens within one answer tend to correlate with large gradient magnitude, entropy alone fails to consistently reflect token importance across different answers, considering the variations in the answer-level reward signals. Based on this observation, we introduce the Gradient Magnitude-based Token Selection (GMTS) method to quantify token importance, which leverages the entropy-gradient connection to approximate gradient-magnitude rankings for token selection. We find that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.

Comment: Findings of the 2026 Conference on Empirical Methods in Natural Language Processing

arXiv abs page · PDF · same-day batch