PaperScope
LIVE · 2026-10-06 05:40 UTC

Prompt Dominance and Asymmetric Verifier Costs: Empirical Ablations of GRPO at 1B Scale on GSM8K

Yi Hou

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04928 v1
Category
Submitted
2026-10-04

Abstract

This paper studies GRPO at 1B scale from both directions: what estimator choices do to the learning signal, and what a degraded reward signal does to what is learned. We train OLMo-2-0425-1B on GSM8K with a from-scratch implementation and measure both sides in controlled sweeps, including a verifier-quality experiment that degrades the training reward and the test-time selector identically. Four results stand out. The prompt is the first-order decision: the zero-shot prompt leaves the base model at 0.08% (its outputs are degenerate continuations, not wrong answers), so almost no group carries a gradient, and training succeeds because the 3-shot prompt reaches 18.3%. At this scale the estimator variants sit within seed noise, with Dr. GRPO ahead on both seeds. In the off-policy regime, clipping is the whole story: training on data without a clipped ratio loses 4-6 points relative to the on-policy reference, while GRPO-style clipping and GSPO recover the loss entirely. Finally, the same weak verifier is far cheaper in RL than in test-time selection: a 10%-flip verifier leaves RL's attainable gain intact (91% and 106% retained across two seeds) where selection retains 57%, and a format-only verifier leaves RL with 16-30% of its gain and selection with essentially nothing. Flip noise acts as an affine transform on the expected reward, and the group-normalized advantage with Adam's rescaling removes it exactly; the residual is a second-order variance effect that the matched-step comparison at a 30% flip rate tests.

Comment: 12 pages, 8 figures. Code, run records, and figure scripts: https://github.com/Helios-YQH/llm-from-scratch

arXiv abs page · PDF · same-day batch