PaperScope
LIVE · 2026-09-29 05:40 UTC

What Does a ProcGen Generalization Gap Measure? Action Rules, Convergence, and the Missing Random Floor

Abhisek Keshari

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32532 v1
Category
Submitted
2026-09-26

Abstract

A generalization gap in reinforcement learning (return on training levels minus return on held-out levels) is usually reported without a reference point. We argue that the missing reference is a measured random floor: the return of a uniform-random policy on the same levels under the same harness. On ProcGen, the floor changes what several standard numbers mean. On identical checkpoints and levels across eight environments, switching between sampled and greedy (argmax) test-time actions moves held-out return in both directions, and greedy evaluation takes three environments to or below the floor: in miner, the sampled policy scores 4.9x the floor on held-out levels while its argmax scores below it. Raw policy entropy places seven of eight environments short of convergence, but 35-65% of that entropy is spread across actions with identical effects; after merging them, one to three remain short, and against the floor only heist has learned nothing that transfers. An audit of twelve prior ProcGen codebases finds that all eleven with held-out evaluation sample test-time actions for their policy-gradient agents, nine by default rather than explicit choice, and six report running in-loop averages rather than evaluating a fixed checkpoint. Applied to our own case study, the same checks grade down a statistically significant encoder effect and rule out a within-encoder train-vs-test CKA statistic. We recommend that every reported gap state its action rule, use a matched and seeded protocol, and report the random floor on both level sets.

arXiv abs page · PDF · same-day batch