TUNDRA: a multilingual corpus of found data for TTS research created with light supervision

TUNDRA: a multilingual corpus of found data for TTS research created with light supervision
复制标题

TUNDRA:在轻度监督下创建的 TTS 研究发现数据的多语言语料库

DOI:
10.21437/interspeech.2013-545
复制
发表时间:
2013
影响因子:
3.4
通讯作者:
Simon King
Simon King
中科院分区:
医学3区
文献类型:
--
作者:
Adriana Stan;O. Watts;Yoshitaka Mamiya;M. Giurgiu;R. Clark;J. Yamagishi;Simon King

文献摘要

被引文献

相似文献

Simple4All Tundra(1.0 版)是标准化多语言语料库的第一个版本,专为不完美或已发现数据的文本到语音研究而设计。该语料库包含来自 14 种语言有声读物的大约 60 小时的语音数据,以及通过轻度监督过程获得的话语级对齐。语料库的未来版本将包括更细粒度的对齐和韵律注释,所有这些都将免费提供。本文概述了迄今为止收集的数据,并详细描述了如何完成此工作,强调了用于编译语料库的最少的特定于语言的知识和手动干预。为了展示其潜在用途,我们使用无监督或轻度监督的方法为所有语言构建了文本转语音系统,论文中也对此进行了简要介绍。索引术语:多语言语料库、轻监督、不完美数据、发现数据、文本转语音、有声读物数据
Simple4All Tundra (version 1.0) is the first release of a standardised multilingual corpus designed for text-to-speech research with imperfect or found data. The corpus consists of approximately 60 hours of speech data from audiobooks in 14 languages, as well as utterance-level alignments obtained with a lightly-supervised process. Future versions of the corpus will include finer-grained alignment and prosodic annotation, all of which will be made freely available. This paper gives a general outline of the data collected so far, as well as a detailed description of how this has been done, emphasizing the minimal language-specific knowledge and manual intervention used to compile the corpus. To demonstrate its potential use, textto-speech systems have been built for all languages using unsupervised or lightly supervised methods, also briefly presented in the paper. Index Terms: multilingual corpus, light supervision, imperfect data, found data, text-to-speech, audiobook data