Myanmar Text-to-Speech System based on Tacotron-2

Myanmar Text-to-Speech System based on Tacotron-2
复制标题

DOI:
10.1109/ictc49870.2020.9289599
复制
发表时间:
2020-10
期刊:
2020 International Conference on Information and Communication Technology Convergence (ICTC)
影响因子:
--
通讯作者:
Yuzana Win;Tomonari Masada
Yuzana Win;Tomonari Masada
中科院分区:
其他
文献类型:
--
作者:
Yuzana Win;Tomonari Masada

文献摘要

被引文献

相似文献

缅甸是东南亚的发展中国家之一,在先进的自然语言处理技术方面仍有许多领域发展不足,文本转语音技术就是其中之一。本文的主要目的是提高缅甸文语转换系统的自然度,使其能够生成类人语音。在本文中,我们应用基于Tacotron-2的神经网络架构,直接从文本序列生成用于语音合成的梅尔频谱图。我们提出的方法由三个步骤组成。在第一步中,我们从大量的新闻文章,小说书籍,日常用法和旅行相关表达中创建了一个包含5 k句子的文本和缅甸文本音频对的语音语料库。我们使用音节分割器和文本规范化器将缅甸语文本分割成一系列字符。在第二步中,我们利用递归序列到序列特征预测网络,将字符嵌入映射到梅尔尺度谱图。在最后一步中,我们使用Griffin-Lim算法将相应的文本转换成缅甸语语音输出。我们将我们提出的方法与基于Tacotron的端到端生成模型进行比较。此外,我们研究了主观评价的两种方法在语音合成中使用平均意见得分(MOS)。实验结果表明,我们提出的方法获得了改善基于Tacotron的语音合成的自然度和可懂度。
Myanmar is one of the developing countries situated in South-East Asia, and there are still many areas that have been under-developed with respect to advanced natural language processing technologies, where text-to-speech is one of them. The main motivation of this paper is to improve the naturalness of Myanmar text-to-speech system that is able to generate human-like speech. In this paper, we apply the neural network architecture based on Tacotron-2 that generates a mel spectrogram for speech synthesis directly from the sequence of text. Our proposed method is composed of three steps. In the first step, we create a speech corpus of 5k sentences of text and audio pair of Myanmar text from a large set of news articles, novel books, daily usages and travel-related expressions. We segment the Myanmar text into a sequence of characters by using a syllable segmenter and text normalizer. In the second step, we utilize the recurrent sequence-to-sequence feature prediction network that maps character embedding to mel-scale spectrograms. In the final step, we use Griffin-Lim algorithm to convert the corresponding text into generate Myanmar speech output. We compare our proposed method with an end-to-end generative model based on Tacotron. Furthermore, we investigate the subjective evaluation for both methods in speech synthesis by using mean opinion score (MOS). The experimental results show that our proposed method obtains an improvement over Tacotron based speech synthesis in terms of naturalness and intelligibility.