PaperScope
LIVE · 2026-09-29 05:40 UTC

The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration

Jin Cui, Xinyue Long, Boran Zhao, Pengju Ren, Hao Dong

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33085 v1
Category
Submitted
2026-09-27

Abstract

Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.

Comment: 23 pages, 5 figures

arXiv abs page · PDF · same-day batch