PaperScope
LIVE · 2026-09-29 05:40 UTC

Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization

Minghao Li, Rui Tan, Ruihang Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32394 v1
Category
Submitted
2026-09-26

Abstract

Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL training run, making sample efficiency a central challenge for reward search on complex control tasks. To address this limitation, we propose an Agentic Reward Black-box Optimization (ARBO) framework, in which an LLM agent builds the search strategy at run time from an evaluation history maintained as its persistent workspace. The evaluation history comprises two components: observations maintained by the evaluation oracle, including candidate scores, per-term training curves, and error tracebacks; and an agent-maintained belief that records diagnoses and intended next steps. The agent queries both with tools and generates the next batch of reward candidates, rather than generating them in a single pass from a fixed prompt. Across four control domains, ARBO achieves gains of 29.9% in manipulation success rate and 192.8% in power-grid score over baseline means under a shared evaluation budget. Ablations examine each component's contribution and sensitivity to backbone choice.

arXiv abs page · PDF · same-day batch