Mellotron: Multispeaker Expressive Voice Synthesis by Conditioning on Rhythm, Pitch and Global Style Tokens

Mellotron: Multispeaker Expressive Voice Synthesis by Conditioning on Rhythm, Pitch and Global Style Tokens
复制标题

DOI:
10.1109/icassp40776.2020.9054556
复制
发表时间:
2019-10
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Rafael Valle;Jason Li;R. Prenger;Bryan Catanzaro
Rafael Valle;Jason Li;R. Prenger;Bryan Catanzaro
中科院分区:
其他
文献类型:
--
作者:
Rafael Valle;Jason Li;R. Prenger;Bryan Catanzaro

文献摘要

被引文献

相似文献

Mellotron是一个基于Tacotron 2 GST的多说话人语音合成模型,它可以在没有情感或歌唱训练数据的情况下使语音表情和歌唱。通过从音频信号或乐谱中明确地调节节奏和连续的音高轮廓,Mellotron能够生成各种风格的语音,从朗读到表达性语音,从缓慢的拖音到说唱,从单调的声音到歌声。与其他方法不同,我们仅使用读取的语音数据来训练Mellotron,而不使用文本和音频之间的对齐。我们使用LJSpeech和LibriTTS数据集来评估我们的模型。我们提供F0帧错误和合成样本,包括来自其他扬声器的风格转移,歌手和训练期间未见过的风格,节奏和音高的程序操作以及合唱合成。
Mellotron is a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data. By explicitly conditioning on rhythm and continuous pitch contours from an audio signal or music score, Mellotron is able to generate speech in a variety of styles ranging from read speech to expressive speech, from slow drawls to rap and from monotonous voice to singing voice. Unlike other methods, we train Mellotron using only read speech data without alignments between text and audio. We evaluate our models using the LJSpeech and LibriTTS datasets. We provide F0 Frame Errors and synthesized samples that include style transfer from other speakers, singers and styles not seen during training, procedural manipulation of rhythm and pitch and choir synthesis.