PaperScope
LIVE · 2026-10-06 05:40 UTC

A Comprehensive Objective Evaluation of Modern Text-to-Speech for Turkish Using Speech Quality Assessment Models

Yunus Emre Ozkose, Alperen Kahraman, Ali Haznedaroglu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06057 v1
Category
Submitted
2026-10-05

Abstract

Modern text-to-speech (TTS) systems can clone a target speaker from a short reference clip or be fine-tuned on a target voice, yet their behaviour on morphologically rich, lower-resource languages such as Turkish remain under-characterised. We present a systematic benchmark of four contemporary systems (Chatterbox, CosyVoice, OmniVoice, and VoxCPM2) evaluated across fine-tuned and zero-shot configurations, contrasted with a conventional VITS baseline and anchored to natural gold speech. Each configuration is scored with eighteen complementary objective metrics spanning learned naturalness predictors (UTMOS v2, DNSMOS-Pro, SCOREQ, WhisQA, AudioBox-PQ, NatScore, SpeechLMScore), intelligibility and signal-quality estimators (SQUIM PESQ/SI-SDR/STOI, Brouhaha), speaker similarity, distributional fidelity (TTSDS) and low-level acoustic descriptors. We further analyse how quality varies with utterance length and quantify long-form temporal consistency through speaker-identity and naturalness drift over chunked utterances. We release our evaluation code to support reproducible TTS evaluation.

Comment: Accepted at 28th International Conference on Speech and Computer (SPECOM 2026). Published in Lecture Notes in Computer Science

Journal: Lecture Notes in Computer Science, SPECOM 2026, Springer, 2026

arXiv abs page · PDF · same-day batch