Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection

Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection
复制标题

DOI:
10.48550/arxiv.2210.14850
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Kentaro Seki;Shinnosuke Takamichi;Takaaki Saeki;H. Saruwatari
Kentaro Seki;Shinnosuke Takamichi;Takaaki Saeki;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
Kentaro Seki;Shinnosuke Takamichi;Takaaki Saeki;H. Saruwatari

文献摘要

被引文献

相似文献

本文提出了一种从暗数据中选择文语转换训练数据的方法。TTS模型通常是在高质量的语音语料库上训练的,这需要花费大量的时间和金钱来收集数据,这使得增加说话人的变化非常具有挑战性。相比之下,有大量数据的可用性是未知的(也称为“暗数据”),例如YouTube视频。为了利用TTS语料库以外的数据,以前的研究已经从语料库中选择语音数据的声学质量的基础上。然而,考虑到已经提出了对数据噪声具有鲁棒性的TTS模型,我们应该根据其作为给定TTS模型的训练数据的重要性来选择数据,而不是语音本身的质量。我们的方法与一个循环的训练和评估选择训练数据的基础上自动预测的质量的合成语音的一个给定的TTS模型。使用YouTube数据的评估结果表明,我们的方法优于传统的基于声学质量的方法。
This paper proposes a method for selecting training data for textto-speech (TTS) synthesis from dark data. TTS models are typically trained on high-quality speech corpora that cost much time and money for data collection, which makes it very challenging to increase speaker variation. In contrast, there is a large amount of data whose availability is unknown (a.k.a, “dark data”), such as YouTube videos. To utilize data other than TTS corpora, previous studies have selected speech data from the corpora on the basis of acoustic quality. However, considering that TTS models robust to data noise have been proposed, we should select data on the basis of its importance as training data to the given TTS model, not the quality of speech itself. Our method with a loop of training and evaluation selects training data on the basis of the automatically predicted quality of synthetic speech of a given TTS model. Results of evaluations using YouTube data reveal that our method outperforms the conventional acoustic-quality-based method.