TrustMed-RL: Long-Horizon Reinforcement Learning for Evidence-Grounded Clinical Diagnosis
Wenxin Zhan, Yizheng Jiao, Haifeng Song, Shuai Xu, Chencheng Pan, Jiayi Feng, Anjie Xie
Abstract
Medical language models can produce correct diagnoses despite incomplete investigations and unsupported reasoning. To support long-horizon, evidence-grounded diagnosis, we introduce \textbf{TrustMed-RL}. Built from PubMed rare-disease cases and over 24,000 manually annotated image panels, it integrates interviews, examinations, testing, specialist consultation, and literature search through state-dependent actions. Our 8B vision--language policy, trained with clinically adapted GiGPO and coverage-adjusted diagnostic rewards, achieves 37.1\% diagnostic accuracy on 2,500 evaluation cases, outperforming all evaluated open-weight baselines and improving over supervised fine-tuning by 12.4 percentage points.. When success additionally requires acquiring at least 50\% of supporting test evidence, TrustMed-RL achieves 32.5\%, exceeding GPT-4o by 6.8 percentage points. Furthermore, it surpasses all evaluated baselines on MTMedDialog and multiple larger 27--32B models on AgentClinic. In physician review of 200 diagnostically accepted test-set trajectories, 83.0\% receive evidential-grounding scores of 4--5 out of 5. Physicians' assessments suggest that these diagnostic trajectories are trustworthy and aligned with human diagnostic reasoning.