TUNDRA: a multilingual corpus of found data for TTS research created with light supervision
TUNDRA: a multilingual corpus of found data for TTS research created with light supervision
复制标题
TUNDRA:在轻度监督下创建的 TTS 研究发现数据的多语言语料库
DOI:
10.21437/interspeech.2013-545
复制
发表时间:
2013
影响因子:
3.4
通讯作者:
Simon King
中科院分区:
文献类型:
--
作者:
Adriana Stan;O. Watts;Yoshitaka Mamiya;M. Giurgiu;R. Clark;J. Yamagishi;Simon King
Simple4All Tundra (version 1.0) is the first release of a standardised multilingual corpus designed for text-to-speech research with imperfect or found data. The corpus consists of approximately 60 hours of speech data from audiobooks in 14 languages, as well as utterance-level alignments obtained with a lightly-supervised process. Future versions of the corpus will include finer-grained alignment and prosodic annotation, all of which will be made freely available. This paper gives a general outline of the data collected so far, as well as a detailed description of how this has been done, emphasizing the minimal language-specific knowledge and manual intervention used to compile the corpus. To demonstrate its potential use, textto-speech systems have been built for all languages using unsupervised or lightly supervised methods, also briefly presented in the paper. Index Terms: multilingual corpus, light supervision, imperfect data, found data, text-to-speech, audiobook data