VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
复制标题
VoxPopuli:用于表征学习、半监督学习和解释的大规模多语言语音语料库
DOI:
10.18653/v1/2021.acl-long.80
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Emmanuel Dupoux
中科院分区:
文献类型:
--
作者:
Changhan Wang;M. Rivière;Ann Lee;Anne Wu;Chaitanya Talnikar;Daniel Haziza;Mary Williamson;J. Pino;Emmanuel Dupoux
We introduce VoxPopuli, a large-scale multilingual corpus providing 400K hours of unlabeled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-supervised learning. VoxPopuli also contains 1.8K hours of transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours. We provide speech recognition (ASR) baselines and validate the versatility of VoxPopuli unlabeled data in semi-supervised ASR and speech-to-text translation under challenging out-of-domain settings. The corpus is available at https://github.com/facebookresearch/voxpopuli.