PaperScope
LIVE · 2026-09-09 05:40 UTC

Risk-Conditioned Fine-Tuning of Large Language Models

Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan zheng

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.08064 v1
Category
Submitted
2026-09-08

Abstract

Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time. In this paper, we propose risk-conditioned RLHF, a framework that trains a single policy that provides a continuous risk-control interface, enabling users to select different degrees of risk aversion without retraining or deploying multiple risk-specific models. Experiments across multiple benchmarks demonstrate that a single risk-conditioned policy can adapt to different risk levels at inference time, enabling more flexible and risk-aware LLM deployment.

Comment: EMNLP 2026 Main

arXiv abs page · PDF · same-day batch