Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining

Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining
复制标题

DOI:
10.48550/arxiv.2301.12596
复制
发表时间:
2023-01
期刊:
--
影响因子:
--
通讯作者:
Takaaki Saeki;Soumi Maiti;Xinjian Li;Shinji Watanabe;Shinnosuke Takamichi;H. Saruwatari
Takaaki Saeki;Soumi Maiti;Xinjian Li;Shinji Watanabe;Shinnosuke Takamichi;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
Takaaki Saeki;Soumi Maiti;Xinjian Li;Shinji Watanabe;Shinnosuke Takamichi;H. Saruwatari

文献摘要

相似文献

虽然神经文本到语音(TTS)已经实现了类似人类的自然合成语音,但由于需要成对的文本和录音室质量的音频数据,多语言TTS系统仅限于资源丰富的语言。本文提出了一种使用目标语言纯文本数据的零触发多语言TTS方法。纯文本数据的使用允许为只有文本资源可用的低资源语言开发TTS系统,使TTS可用于数千种语言。受多语言语言模型强大的跨语言可移植性的启发,我们的框架首先使用多语言纯文本数据进行掩码语言模型预训练。然后,我们以监督的方式使用配对数据训练这个模型,同时冻结语言感知的嵌入层。这允许即使对于不包括在配对数据中但存在于纯文本数据中的语言也进行推断。评估结果表明,高度可理解的零拍TTS的字符错误率小于12%的一个看不见的语言。
While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method for zero-shot multilingual TTS using text-only data for the target language. The use of text-only data allows the development of TTS systems for low-resource languages for which only textual resources are available, making TTS accessible to thousands of languages. Inspired by the strong cross-lingual transferability of multilingual language models, our framework first performs masked language model pretraining with multilingual text-only data. Then we train this model with a paired data in a supervised manner, while freezing a language-aware embedding layer. This allows inference even for languages not included in the paired data but present in the text-only data. Evaluation results demonstrate highly intelligible zero-shot TTS with a character error rate of less than 12% for an unseen language.