VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
复制标题

VoxPopuli:用于表征学习、半监督学习和解释的大规模多语言语音语料库

DOI:
10.18653/v1/2021.acl-long.80
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Emmanuel Dupoux
Emmanuel Dupoux
中科院分区:
--
文献类型:
--
作者:
Changhan Wang;M. Rivière;Ann Lee;Anne Wu;Chaitanya Talnikar;Daniel Haziza;Mary Williamson;J. Pino;Emmanuel Dupoux

文献摘要

被引文献

相似文献

我们介绍了一个大规模的多语种语料库voxopi,该语料库提供了23种语言的40万小时的未标记语音数据。这是迄今为止非监督表示学习和半监督学习的最大公开数据。VOXPUPI还包含15种语言的1.8K小时转录演讲及其15种目标语言的对准口头口译,总计17.3K小时。我们提供了语音识别(ASR)基线,并在具有挑战性的域外环境下验证了半监督ASR和语音到文本翻译中Voxopi未标记数据的通用性。语料库可在https://github.com/facebookresearch/voxpopuli.上找到
We introduce VoxPopuli, a large-scale multilingual corpus providing 400K hours of unlabeled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-supervised learning. VoxPopuli also contains 1.8K hours of transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours. We provide speech recognition (ASR) baselines and validate the versatility of VoxPopuli unlabeled data in semi-supervised ASR and speech-to-text translation under challenging out-of-domain settings. The corpus is available at https://github.com/facebookresearch/voxpopuli.