PaperScope
LIVE · 2026-09-29 05:40 UTC

CompassPlay: Rewarding the Proposer for Where It Moves the Solver

Sophia Xiao Pu, Ximeng Sun, Jiang Liu, Jialian Wu, Emad Barsoum, Zicheng Liu, William Yang Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32228 v1
Category
Submitted
2026-09-26

Abstract

In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks whose solver loss gradients align with those of reference tasks representing the target capabilities. It draws on a first-order approximation to learning progress and scores each eligible task without additional solver training. Our experiments show gains in performance and training efficiency. In coding self-play with Qwen2.5-Coder-7B, a small external reference set guides task generation. CompassPlay improves average accuracy over AZR's difficulty reward by 1.5 percentage points on in-domain coding and 2.7 on out-of-domain mathematics. In Lean4 theorem proving, CompassPlay matches the difficulty baseline's 150-iteration cumulative coverage with 40\% fewer GPU-hours.

arXiv abs page · PDF · same-day batch