PaperScope
LIVE · 2026-09-22 05:40 UTC

HaikuS2S: A Cascaded System For Responding In Verse

Devangi Sharma, Sophia Judicke, Glenda Tan, Conrad Schaumburg, Shinji Watanabe

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23951 v1
Category
Submitted
2026-09-20

Abstract

Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation. However, these models do not model haiku's 5-7-5 syllable structure or line-ending pauses. We present a cascaded system, HaikuS2S, combining ASR (automatic speech recognition), LLM (large language model)-generated haiku, and TTS fine-tuning on both prose and custom haiku datasets. Our evaluation focuses on emotion similarity, speech quality, and prosody alignment. In our experiments, we see that our prosody and tonal alignment improve significantly with our fine-tuned systems, particularly the one trained on both general poetry and haiku. We also see that we maintain similar emotion similarity scores across all systems.

Comment: Accepted to SLT 2026, Demo Track. 5 pages, 5 figures

arXiv abs page · PDF · same-day batch