PaperScope
LIVE · 2026-09-29 05:40 UTC

$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training

Nan Qiao, Yebin Yang, Weinong Wang, Shuning Wang, Shangpin Peng, Fengyuan Lu, Xinming Wang, Zhehan Kan, Ruixu Zhang, Songyang Zhang, Sheng Yue, Yonglong Tian, Ju Ren

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32791 v1
Category
Submitted
2026-09-26

Abstract

Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.

arXiv abs page · PDF · same-day batch