PaperScope
LIVE · 2026-10-06 05:40 UTC

Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking

Ali Vosoughi, Akhil Kasturi, Chenliang Xu, Axel Wismueller

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05425 v1
Category
Submitted
2026-10-04

Abstract

Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.

Comment: 40 pages, 7 figures, 17 tables (main text and references pp. 1-19; Supplementary Information as an appendix, pp. 20-40). Submitted to npj Digital Medicine. Code: https://github.com/ali-vosoughi/VerifyGRPO-Rad

arXiv abs page · PDF · same-day batch