Utterance Selection Techniques for TTS Systems Using Found Speech
Utterance Selection Techniques for TTS Systems Using Found Speech
复制标题
使用找到的语音的 TTS 系统的话语选择技术
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
A. Black
中科院分区:
文献类型:
--
作者:
P. Baljekar;A. Black
The goal in this paper is to investigate data selection techniques for found speech. Found speech unlike clean, phonetically-balanced datasets recorded specifically for synthesis contain a lot of noise which might not get labeled well and it might contain utterances with varying channel conditions. These channel variations and other noise distortions might sometimes be useful in terms of adding diverse data to our training set, however in other cases it might be detrimental to the system. The approach outlined in this work investigates various metrics to detect noisy data which degrade the performance of the system on a held-out test set. We assume a seed set of 100 utterances to which we then incrementally add in a fixed set of utterances and find which metrics can capture the misaligned and noisy data. We report results on three datasets, an artificially degraded set of clean speech, a single speaker database of found speech and a multi - speaker database of found speech. All of our experiments are carried out on male speakers. We also show comparable results are obtained on a female multi-speaker corpus.