PaperScope
LIVE · 2026-09-29 05:40 UTC

From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS

Kangxiang Xia, Xinfa Zhu, HangRui Hu, Kexin Huang, Wenjie Tian, Ziyue Jiang, Bingshen Mu, Jingbin Hu, Ting He, Lei Xie, Jin Xu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33362 v1
Submitted
2026-09-27

Abstract

Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.

arXiv abs page · PDF · same-day batch