PaperScope
LIVE · 2026-10-07 05:40 UTC

ARIA: Audio-Driven Melody-Tone Relation Modeling for Cantonese Lyric Authoring

Shengyu Li, Jinting Wang, Li Liu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.07902 v1
Category
Submitted
2026-10-06

Abstract

Cantonese lyric writing requires close alignment between lexical tones and melodic pitch. Existing melody-guided lyric generation methods typically rely on symbolic melody to generate lyrics. However, in real songwriting scenarios, melodies are often expressed as raw singing audio or hummed recordings, where pitch is implicit, noisy, and unstructured, making these methods difficult to apply directly. To address this limitation, we propose ARIA, a two-stage audio-driven melody-tone relation modeling framework for Cantonese lyric authoring that generates Cantonese lyrics from singing recordings with provided character-level timestamps. Specifically, we first design a Tri-Stream Relation-Aware Tone Estimator (TRATE) to predict 0243 sequences from timestamped singing audio by modeling multi-stream acoustic cues and relational tonal structure. We then propose a Decoupled Retrieval-Augmented Tone-Conditioned Lyric Generator (DRA-TCLG) to generate fluent lyrics conditioned on predicted tonal plans with retrieval-enhanced lexical guidance. Moreover, we construct a large-scale aligned audio-Jyutping-0243 dataset from real Cantonese singing recordings to support this new task. Experimental results demonstrate that ARIA achieves strong performance in both 0243 prediction and tone-consistent lyric generation, validating the effectiveness of the proposed framework.

Comment: Accepted for publication in Findings of EMNLP 2026. 24 pages, including references and appendices. Author-prepared version

arXiv abs page · PDF · same-day batch