SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment
De Jiang, Shuo Zhang, Kehong Yuan, Hongen Liao
Abstract
Developmental screening videos show how children perform specific behaviors, but clinical records usually contain outcomes rather than descriptions of what happened. We present SemReward-VL, which learns to describe item-specific behavior from these outcomes. A vision-language model generates a description, and a frozen language model scores its agreement with the clinical outcome, relevance to the item, abstention on unrelated video-item pairs, and clarity. Group relative policy optimization (GRPO) updates LoRA adapters using this semantic reward. On 13,379 videos covering 41 items, the method improves aggregate accuracy and the number of items with recall above 0.5. Errors remain in temporal direction, duration, and age-specific interpretations of behavior.