Characteristics of Text-to-Speech and Other Corpora

Characteristics of Text-to-Speech and Other Corpora
复制标题

DOI:
10.21437/speechprosody.2018-140
复制
发表时间:
2018-06
期刊:
Speech Prosody 2018
影响因子:
--
通讯作者:
Erica Cooper;E. Li;Julia Hirschberg
Erica Cooper;E. Li;Julia Hirschberg
中科院分区:
其他
文献类型:
--
作者:
Erica Cooper;E. Li;Julia Hirschberg

文献摘要

被引文献

相似文献

大量的TTS语料库存在于为高资源语言(如普通话、英语和日语)创建的商业系统中。为这些语料库录制的扬声器通常被指示保持恒定的f0、能量和说话速率,并在理想的声学环境中录制,产生干净、一致的音频。我们一直在根据为其他目的(例如培训ASR系统)或网络上可用(例如新闻广播、有声读物)收集的“发现”数据开发TTS系统,以针对低资源语言(LRL)制作TTS系统,目前没有昂贵的商业系统。这项研究调查了传统TTS扬声器是否表现出显着更少的变化和更好的说话特点比扬声器在发现流派。通过对所发现的语音类型的f_0、能量、语速、清晰度、NHR、抖动和shimmer的特征进行分析,并与传统的TTS语料库进行比较,我们发现TTS录音确实具有较低的平均音高、能量标准差、语速和清晰度水平,以及较低的shimmer和NHR的平均和标准差;在许多方面,这些与一些已发现的流派非常相似。通过识别相似性和差异性,我们能够识别选择发现的数据来构建LRL TTS系统的客观方法。
Extensive TTS corpora exist for commercial systems created for high-resource languages such as Mandarin, English, and Japanese. Speakers recorded for these corpora are typically instructed to maintain constant f0, energy, and speaking rate and are recorded in ideal acoustic environments, producing clean, consistent audio. We have been developing TTS systems from “found” data collected for other purposes (e.g. training ASR systems) or available on the web (e.g. news broadcasts, audiobooks) to produce TTS systems for low-resource languages (LRLs) which do not currently have expensive, commercial systems. This study investigates whether traditional TTS speakers do exhibit significantly less variation and better speaking characteristics than speakers in found genres. By examining characteristics of f0, energy, speaking rate, articulation, NHR, jitter, and shimmer in found genres and comparing these to traditional TTS corpora, We have found that TTS recordings are indeed characterized by low mean pitch, standard deviation of energy, speaking rate, and level of articulation, and low mean and standard deviations of shimmer and NHR; in a number of respects these are quite similar to some found genres. By identifying similarities and differences, we are able to identify objective methods for selecting found data to build TTS systems for LRLs.