Utterance Selection Techniques for TTS Systems Using Found Speech

Utterance Selection Techniques for TTS Systems Using Found Speech
复制标题

使用找到的语音的 TTS 系统的话语选择技术

DOI:
--
复制
发表时间:
2016
期刊:
Speech Synthesis Workshop
影响因子:
--
通讯作者:
A. Black
A. Black
中科院分区:
--
文献类型:
--
作者:
P. Baljekar;A. Black

文献摘要

被引文献

相似文献

本文的目标是调查数据选择技术,发现语音。与专门为合成而记录的干净、语音平衡的数据集不同,发现的语音包含大量噪声,这些噪声可能无法很好地标记,并且可能包含具有不同信道条件的话语。这些信道变化和其他噪声失真有时可能有助于将不同的数据添加到我们的训练集,但在其他情况下,它可能对系统有害。在这项工作中概述的方法调查各种指标来检测噪声数据,降低了系统的性能,在一个举行了测试集。我们假设一个100个话语的种子集,然后我们递增地添加到一个固定的话语集,并找到哪些度量可以捕获未对齐和噪声数据。我们报告了三个数据集的结果,一个阿尔蒂正式退化的干净语音集,一个单一的说话人数据库的发现语音和多说话人数据库的发现语音。我们所有的实验都是在男性说话者身上进行的。我们还显示了可比的结果上获得的女性多说话人语料库。
The goal in this paper is to investigate data selection techniques for found speech. Found speech unlike clean, phonetically-balanced datasets recorded specifically for synthesis contain a lot of noise which might not get labeled well and it might contain utterances with varying channel conditions. These channel variations and other noise distortions might sometimes be useful in terms of adding diverse data to our training set, however in other cases it might be detrimental to the system. The approach outlined in this work investigates various metrics to detect noisy data which degrade the performance of the system on a held-out test set. We assume a seed set of 100 utterances to which we then incrementally add in a fixed set of utterances and find which metrics can capture the misaligned and noisy data. We report results on three datasets, an artificially degraded set of clean speech, a single speaker database of found speech and a multi - speaker database of found speech. All of our experiments are carried out on male speakers. We also show comparable results are obtained on a female multi-speaker corpus.