Perceptual Quality Dimensions of Text-to-Speech Systems
Perceptual Quality Dimensions of Text-to-Speech Systems
复制标题
文本转语音系统的感知质量维度
DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
U. Heute
中科院分区:
文献类型:
--
作者:
Florian Hinterleitner;Sebastian Möller;C. Norrenbrock;U. Heute
The aim of this paper is to analyze the perceptual quality dimensions of state-of-the-art text-to-speech systems (TTS). Therefore, several pretests were conducted to determine a suitable set of attribute scales. The resulting 16 scales were used in a semantic differential on a diverse database containing 16 different TTS systems. A subsequent multidimensional analysis (Principal Axis Factor analysis with Promax rotation) resulted in three underlying quality dimensions. They were labeled naturalness, disturbances, and temporal distortions. A mapping of these factors onto the perceived overall quality revealed that naturalness contributes the most to the quality of TTS signals. Index Terms: speech synthesis, quality dimensions, multidimensional analysis