Perceptual Quality Dimensions of Text-to-Speech Systems

Perceptual Quality Dimensions of Text-to-Speech Systems
复制标题

文本转语音系统的感知质量维度

DOI:
--
复制
发表时间:
2011
期刊:
Interspeech
影响因子:
--
通讯作者:
U. Heute
U. Heute
中科院分区:
--
文献类型:
--
作者:
Florian Hinterleitner;Sebastian Möller;C. Norrenbrock;U. Heute

文献摘要

被引文献

相似文献

本文的目的是分析国家的最先进的文语转换系统(TTS)的感知质量维度。因此,进行了几次预测试,以确定一组合适的属性量表。由此产生的16个尺度被用于在一个不同的数据库,包含16个不同的TTS系统的语义差异。随后的多维分析(主轴因子分析与Promax旋转)导致三个基本的质量维度。它们被贴上了自然、干扰和时间扭曲的标签。将这些因素映射到感知的整体质量上,发现自然度对TTS信号的质量贡献最大。索引术语:语音合成,质量维度,多维分析
The aim of this paper is to analyze the perceptual quality dimensions of state-of-the-art text-to-speech systems (TTS). Therefore, several pretests were conducted to determine a suitable set of attribute scales. The resulting 16 scales were used in a semantic differential on a diverse database containing 16 different TTS systems. A subsequent multidimensional analysis (Principal Axis Factor analysis with Promax rotation) resulted in three underlying quality dimensions. They were labeled naturalness, disturbances, and temporal distortions. A mapping of these factors onto the perceived overall quality revealed that naturalness contributes the most to the quality of TTS signals. Index Terms: speech synthesis, quality dimensions, multidimensional analysis