Investigation of Enhanced Tacotron Text-to-speech Synthesis Systems with Self-attention for Pitch Accent Language

Investigation of Enhanced Tacotron Text-to-speech Synthesis Systems with Self-attention for Pitch Accent Language
复制标题

DOI:
10.1109/icassp.2019.8682353
复制
发表时间:
2018-10
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Yusuke Yasuda;Xin Wang;Shinji Takaki;J. Yamagishi
Yusuke Yasuda;Xin Wang;Shinji Takaki;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Yusuke Yasuda;Xin Wang;Shinji Takaki;J. Yamagishi

文献摘要

被引文献

相似文献

端到端语音合成是一种很有前途的方法,它直接将原始文本转换为语音。尽管Tacotron 2在英语的自然性方面优于经典管道系统,但其对其他语言的适用性仍然是未知的。日语可能是实现端到端语音合成最困难的语言之一,主要是由于其字符多样性和音调口音。因此,最先进的系统仍然基于传统的管道框架,需要单独的文本分析器和持续时间模型。为了实现端到端的日语语音合成,我们将Tacotron扩展到具有自我注意力的系统,以捕获与音高口音相关的长期依赖关系,并将其音频质量与各种条件下的经典管道系统进行比较,以显示其优缺点。在大规模听力测试中,我们研究了重音类型标签的存在,使用力或预测对齐的影响,以及用作Wavenet声码器的局部条件参数的声学特征。我们的研究结果表明,虽然所提出的系统仍然不匹配的顶级流水线系统的质量为日本,我们显示了重要的垫脚石端到端的日语语音合成。
End-to-end speech synthesis is a promising approach that directly converts raw text to speech. Although it was shown that Tacotron2 outperforms classical pipeline systems with regards to naturalness in English, its applicability to other languages is still unknown. Japanese could be one of the most difficult languages for which to achieve end-to-end speech synthesis, largely due to its character diversity and pitch accents. Therefore, state-of-the-art systems are still based on a traditional pipeline framework that requires a separate text analyzer and duration model. Towards end-to-end Japanese speech synthesis, we extend Tacotron to systems with self-attention to capture long-term dependencies related to pitch accents and compare their audio quality with classical pipeline systems under various conditions to show their pros and cons. In a large-scale listening test, we investigated the impacts of the presence of accentual-type labels, the use of force or predicted alignments, and acoustic features used as local condition parameters of the Wavenet vocoder. Our results reveal that although the proposed systems still do not match the quality of a top-line pipeline system for Japanese, we show important stepping stones towards end-to-end Japanese speech synthesis.