PaperScope
LIVE · 2026-09-29 05:40 UTC

Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers

Qiangqiang He, Jin Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32674 v1
Category
Submitted
2026-09-26

Abstract

On-policy reasoning models can learn from task rewards or teacher signals, but these sources differ in form and can favor conflicting updates, leaving unclear which should guide a given reasoning action. We introduce \textbf{Expected Reasoning-Step Return (ERSR)}, which treats semantic reasoning steps as macro-actions and uses Monte Carlo student-policy rollouts to estimate the expected final task reward of student-generated and teacher-proposed actions in a common return space for step-level comparison. ERSR analysis reveals an outcome-dependent asymmetry: student actions are more beneficial than teacher replacements on successful trajectories, whereas teacher replacements become more beneficial on failed trajectories. We further show that student answer-probe gains track student-step ERSR utility and distinguish beneficial from harmful reasoning steps. Based on these findings, we propose \textbf{Return-Referenced On-Policy Learning (R$^2$OPL)}, which reinforces student reasoning on successful trajectories and distills teacher signals on failed ones, while using group success rate for difficulty scaling and student-probe gains for step-level modulation. Experiments across reasoning benchmarks and teacher--student configurations show that R$^2$OPL consistently outperforms strong baselines. ERSR training dynamics further show that R$^2$OPL jointly exploits substantial utility from both reward- and teacher-side signals, whereas existing hybrids often leave substantial residual utility in one branch.

Comment: 32 pages, 10 figures

arXiv abs page · PDF · same-day batch