PaperScope
LIVE · 2026-09-29 05:40 UTC

SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale

Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai, Rei Kawakami

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35225 v1
Category
Submitted
2026-09-28

Abstract

Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.

Comment: Accepted by EMNLP 2026 Findings

arXiv abs page · PDF · same-day batch