Investigation of Using Continuous Representation of Various Linguistic Units in Neural Network Based Text-to-Speech Synthesis

Investigation of Using Continuous Representation of Various Linguistic Units in Neural Network Based Text-to-Speech Synthesis
复制标题

DOI:
10.1587/transinf.2016slp0011
复制
发表时间:
2016-10
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Xin Wang;Shinji Takaki;Junichi Yamagishi
Xin Wang;Shinji Takaki;Junichi Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Xin Wang;Shinji Takaki;Junichi Yamagishi

文献摘要

相似文献

在不了解目标语言的情况下构建高质量的文本到语音(TTS)系统是一个重要但具有挑战性的研究课题。在这种TTS系统中,寻找既有效又易于获取的输入文本表示是至关重要的。最近,被称为“词嵌入”的原始词输入的连续表示已经成功地应用于各种自然语言处理任务中。它也被用作TTS系统的基于神经网络的声学模型的附加或替代语言输入特征。在本文中,我们进一步研究了使用这种嵌入技术来表示基于循环和前馈神经网络的声学模型的音素、音节和短语。实验结果表明,当这些连续表征作为附加成分或替代传统韵律上下文输入声学模型时,大多数连续表征都不能显著提高系统的性能。然而,主观评价表明,短语的连续表示与韵律上下文相结合,作为基于前馈神经网络的声学模型的输入,可以取得显著的改善。
SUMMARY Building high-quality text-to-speech (TTS) systems without expert knowledge of the target language and / or time-consuming manual annotation of speech and text data is an important yet challenging research topic. In this kind of TTS system, it is vital to find representation of the input text that is both e ff ective and easy to acquire. Recently, the continuous representation of raw word inputs, called “word embedding”, has been successfully used in various natural language processing tasks. It has also been used as the additional or alternative linguistic input features to a neural-network-based acoustic model for TTS systems. In this paper, we further investigate the use of this embedding technique to represent phonemes, syllables and phrases for the acoustic model based on the recurrent and feed-forward neural network. Results of the experiments show that most of these continuous representations cannot significantly improve the system’s performance when they are fed into the acoustic model either as additional component or as a replacement of the conventional prosodic context. However, subjective evaluation shows that the continuous representation of phrases can achieve significant improvement when it is combined with the prosodic context as input to the acoustic model based on the feed-forward neural network.