PaperScope
LIVE · 2026-10-06 05:40 UTC

A Bird's-Eye View of Iterative Reward Design

Logan Mondal Bhamidipaty, Lauren Robson, Linda Petrini, Shengrui Lyu, Kamal Ndousse

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04364 v1
Category
Submitted
2026-10-03

Abstract

Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback. However, these methods are often hard to compare because they differ in implementation details, feedback assumptions, and evaluation environments. To address this, we introduce a Benchmark for Iterative Reward Design (BIRD) that expresses existing methods in a unified configuration and evaluation space. This lets us compare algorithms directly, ablate individual design choices, and prototype new components under matched feedback conditions and policy-training budgets. Across MuJoCo, Meta-World, Assistax, and HumanoidBench, we identify a small set of simple design choices that consistently improve performance. Combining these choices yields significantly better performance than the evaluated methods from prior work. Our results highlight the strength of simple baselines and motivate further study of when additional algorithmic complexity improves iterative reward design. Code is available at https://github.com/safety-research/bird.

Comment: 30 pages

arXiv abs page · PDF · same-day batch