PaperScope
LIVE · 2026-09-09 05:40 UTC

Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation

Wenbo Zhang, Wenzhuo Zhou, Hengrui Cai, Zhengling Qi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06893 v1
Category
Submitted
2026-09-07

Abstract

Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterative procedure, is a key factor behind its superiority over offline methods. This insight leads us to answer the second question affirmatively and propose Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm. DP3O first learns an explicit preference model using a helper class of LLMs and then distills its knowledge into policy optimization. Theoretically, we show that explicit preference modeling admits better estimation error control than implicit formulations, and that DP3O achieves a tighter generalization bound than hard-label DPO through variance reduction. Empirically, we evaluate DP3O on a wide range of chat-based and downstream tasks and show that it outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time by about $42\%$, demonstrating both its effectiveness and efficiency.

Comment: Accepted by Transactions on Machine Learning Research (TMLR)

arXiv abs page · PDF · same-day batch