PaperScope
LIVE · 2026-09-03 05:40 UTC

Dense Process Supervision for Search Agents via Fact Utility Estimation

Rongzhi Zhu, Xiangyu Liu, Yi Liu, Shuo Zhang, Ruirui Zhang, Rui Wu, Tao Jiang, Zequn Sun, Wenhao Xu, Wei Hu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.00833 v1
Category
Submitted
2026-09-01

Abstract

Reinforcement learning (RL) for search agents typically relies on outcome rewards. However, it often fails to achieve effective credit assignment, due to the unclear value of intermediate steps. It is hard to separate their contributions from the final result. In this paper, we propose a dense process supervision method based on fact utility estimation, which models the reasoning process as the accumulation of discrete evidence facts. We first extract structured facts from raw observations and organize them into an explicit fact store. To support credit assignment, we then cluster semantically equivalent facts and infer the posterior utility of each fact cluster using Bayesian estimation over group rollouts. Finally, we convert the estimated fact utilities into dense step-level rewards to guide RL training. Experiments on seven single-hop and multi-hop QA benchmarks show that our method consistently outperforms existing baselines. Ablation studies validate clear relative improvements on multi-hop QA compared to outcome reward-only training.

Comment: Accepted in the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

arXiv abs page · PDF · same-day batch