PaperScope
LIVE · 2026-09-09 05:40 UTC

SignRefine: Adapting Foundational Video Models for Sign Language Generation

Anton Pelykh, Edward Fish, Ozge Mercanoglu Sincan, Richard Bowden

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.08496 v1
Category
Submitted
2026-09-08

Abstract

Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions. Our approach builds on a pretrained video diffusion transformer and introduces local adapters with spatial grounding to selectively refine hand and face regions, steering the strong base model's prior toward accurate articulation. To enable this work and support broader sign language research, we present NVSign, a large-scale dataset of video content natively produced in sign language, offering diverse signer appearances, environments, and natural conversational settings. Trained on this data, our model shows up to 30% improvement in hand pose precision metrics over the strongest baseline and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.

arXiv abs page · PDF · same-day batch