Towards Signal-Based Instrumental Quality Diagnosis for Text-to-Speech Systems
Towards Signal-Based Instrumental Quality Diagnosis for Text-to-Speech Systems
复制标题
面向文本转语音系统的基于信号的仪器质量诊断
DOI:
--
复制
发表时间:
2008
影响因子:
3.9
通讯作者:
Sebastian Möller
中科院分区:
文献类型:
--
作者:
T. Falk;Sebastian Möller
In this letter, the first steps toward the development of a signal-based instrumental quality measure for text-to-speech (TTS) systems are described. Hidden Markov models (HMM), trained on naturally-produced speech, serve as artificial text- and speaker-independent reference models against which synthesized speech signals are assessed. A normalized log-likelihood measure, computed between perceptual features extracted from synthesized speech and a gender-dependent HMM reference model, is proposed and shown to be a reliable parameter for multidimensional TTS quality diagnosis. Experiments with subjectively scored synthesized speech data show that the proposed measure attains promising estimation performance for quality dimensions labeled overall impression, listening effort, naturalness, continuity/fluency, and acceptance.