Neural Speech Synthesis with Transformer Network

Neural Speech Synthesis with Transformer Network
复制标题

DOI:
10.1609/aaai.v33i01.33016706
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Naihan Li;Shujie Liu;Yanqing Liu;Sheng Zhao;Ming Liu-
Naihan Li;Shujie Liu;Yanqing Liu;Sheng Zhao;Ming Liu-
中科院分区:
其他
文献类型:
--
作者:
Naihan Li;Shujie Liu;Yanqing Liu;Sheng Zhao;Ming Liu-

文献摘要

被引文献

相似文献

虽然端到端神经文本到语音(TTS)方法(如Tacotron 2)被提出并实现了最先进的性能,但它们仍然存在两个问题:1)在训练和推理过程中效率低下; 2)难以使用当前的递归神经网络(RNN)对长依赖性进行建模。受Transformer网络在神经机器翻译(NMT)中的成功启发,本文引入并调整了多头注意力机制,以取代RNN结构以及Tacotron 2中的原始注意力机制。在编码器和解码器中,利用多头自注意并行构造隐藏状态,提高了训练效率。同时,通过自注意机制将不同时刻的任意两个输入直接连接起来,有效地解决了长距离依赖问题。使用音素序列作为输入,我们的Transformer TTS网络生成mel频谱图,然后由WaveNet声码器输出最终的音频结果。实验进行了测试我们的新网络的效率和性能。在效率方面,我们的Transformer TTS网络与Tacotron 2相比,可以加快训练速度约4.25倍。对于性能,严格的人体测试表明,我们提出的模型达到了最先进的性能(优于Tacotron 2,差距为0.048),并且非常接近人类质量(MOS中为4.39 vs 4.44)。
Although end-to-end neural text-to-speech (TTS) methods (such as Tacotron2) are proposed and achieve state-of-theart performance, they still suffer from two problems: 1) low efficiency during training and inference; 2) hard to model long dependency using current recurrent neural networks (RNNs). Inspired by the success of Transformer network in neural machine translation (NMT), in this paper, we introduce and adapt the multi-head attention mechanism to replace the RNN structures and also the original attention mechanism in Tacotron2. With the help of multi-head self-attention, the hidden states in the encoder and decoder are constructed in parallel, which improves training efficiency. Meanwhile, any two inputs at different times are connected directly by a self-attention mechanism, which solves the long range dependency problem effectively. Using phoneme sequences as input, our Transformer TTS network generates mel spectrograms, followed by a WaveNet vocoder to output the final audio results. Experiments are conducted to test the efficiency and performance of our new network. For the efficiency, our Transformer TTS network can speed up the training about 4.25 times faster compared with Tacotron2. For the performance, rigorous human tests show that our proposed model achieves state-of-the-art performance (outperforms Tacotron2 with a gap of 0.048) and is very close to human quality (4.39 vs 4.44 in MOS).