PaperScope
LIVE · 2026-09-29 05:40 UTC

Agentic Multi-Turn Reasoning: A Fairness Approach

Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill, Jackson Cothren, Marios Savvides, Khoa Luu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33323 v1
Category
Submitted
2026-09-27

Abstract

Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $Φ$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.

Comment: Accepted to NeurIPS'26

arXiv abs page · PDF · same-day batch