MuST-C: a Multilingual Speech Translation Corpus
MuST-C: a Multilingual Speech Translation Corpus
复制标题
DOI:
10.18653/v1/n19-1202
复制
发表时间:
2019-06
期刊:
影响因子:
--
通讯作者:
Mattia Antonino Di Gangi;R. Cattoni;L. Bentivogli;Matteo Negri;Marco Turchi
中科院分区:
文献类型:
--
作者:
Mattia Antonino Di Gangi;R. Cattoni;L. Bentivogli;Matteo Negri;Marco Turchi
Current research on spoken language translation (SLT) has to confront with the scarcity of sizeable and publicly available training corpora. This problem hinders the adoption of neural end-to-end approaches, which represent the state of the art in the two parent tasks of SLT: automatic speech recognition and machine translation. To fill this gap, we created MuST-C, a multilingual speech translation corpus whose size and quality will facilitate the training of end-to-end systems for SLT from English into 8 languages. For each target language, MuST-C comprises at least 385 hours of audio recordings from English TED Talks, which are automatically aligned at the sentence level with their manual transcriptions and translations. Together with a description of the corpus creation methodology (scalable to add new data and cover new languages), we provide an empirical verification of its quality and SLT results computed with a state-of-the-art approach on each language direction.