A comparison of open-source segmentation architectures for dealing with imperfect data from the media in speech synthesis

A comparison of open-source segmentation architectures for dealing with imperfect data from the media in speech synthesis
复制标题

用于处理语音合成中来自媒体的不完美数据的开源分段架构的比较

DOI:
10.21437/interspeech.2014-515
复制
发表时间:
2014
影响因子:
3.4
通讯作者:
Simon King
Simon King
中科院分区:
医学3区
文献类型:
--
作者:
A. Gallardo;J. Montero;Simon King

文献摘要

被引文献

相似文献

传统的文本到语音(TTS)系统是使用特别设计的非表达脚本记录开发的。为了在Simple4All项目中开发新一代富有表现力的TTS系统,应该使用来自媒体的真实录音来训练具有全新说话风格的新声音。然而,为了处理这种更自然的材料,新系统必须能够处理不完美的数据(多扬声器录音,背景和前景音乐和噪音),过滤掉低质量的音频片段并创建单扬声器集群。在本文中,我们比较了几种将说话人分割与音乐和噪声检测相结合的体系结构,这些体系结构提高了分割的精度和整体质量。
Traditional Text-To-Speech (TTS) systems have been developed using especially-designed non-expressive scripted recordings. In order to develop a new generation of expressive TTS systems in the Simple4All project, real recordings from the media should be used for training new voices with a whole new range of speaking styles. However, for processing this more spontaneous material, the new systems must be able to deal with imperfect data (multi-speaker recordings, background and foreground music and noise), filtering out low-quality audio segments and creating mono-speaker clusters. In this paper we compare several architectures for combining speaker diarization and music and noise detection which improve the precision and overall quality of the segmentation.